CORTEXA
← Browse
crossrefAI2026-03-02Cited by 0

The Development of a Large Language Model-Powered Chatbot to Advance Fairness in Machine Learning

Pedro Henrique Ribeiro Santiago, Xiangqun Ju, Xavier Vasquez, Heidi Shen, Lisa Jamieson, Hawazin W. Elani

Background: Machine learning (ML) has been widely adopted in decision-making, making fairness a central ethical and scientific priority. We developed the Themis chatbot, a Large Language Model (LLM) system designed to explain concepts of ML fairness in an accessible, conversational format. Methods: The development followed four stages: (1) curating a document corpus of 286 peer-reviewed publications on ML fairness; (2) development of Themis by combining a modern LLM (OpenAI’s GPT-4o) with Retrieval Augmented Generation (RAG); (3) creation of a 340-item benchmark dataset, the FairnessQA; and (4) evaluating performance against state-of-the-art non-augmented LLMs (DeepSeek R1, GPT-4o, GPT-5, and Grok 3). Results: For the multiple-choice questions, Themis achieved an accuracy of 96.7%, outperforming DeepSeek R1 (90.0%), GPT-4o (89.3%), GPT-5 (92.0%), and Grok 3 (86.7%), and the overall difference was statistically significant (χ2(4) = 10.1, p = 0.038). In the closed-ended questions, Themis achieved the highest accuracy (96.7%), while competing models ranged from 78.0% to 84.0%, and the overall difference was significant (χ2(4) = 23.9, p < 0.001). In the open-ended questions, Themis achieved the highest mean scores for correctness (M = 4.62), completeness (M = 4.59), and usefulness (M = 4.56), and differences were statistically significant (correctness: F(4, 195) = 20.91, p < 0.001; completeness: F(4, 195) = 7.76, p < 0.001; usefulness: F(4, 195) = 2.90, p < 0.001). By consolidating scattered research into an interactive assistant, Themis makes fairness concepts more accessible to educators, researchers, and policymakers. This work demonstrates that retrieval-augmented systems can enhance the public understanding of machine learning fairness at scale.

View free PDFSource page

Related papers

crossrefAI2026-07-12

eGFR-AI: A Stacked Machine-Learning Model for Early Postoperative Kidney Function Prediction—A Pilot Study

Eva Brenner, Luka Bulić, Vilena Vrbanović Mijatović

Background: Postoperative kidney dysfunction is a common and serious complication in surgical patients. Kidney function is typically assessed using the estimated glomerular filtration rate (eGFR), most often calculated with the CKD-EPI equation based on serum creatinine. While se…

View free PDFSource page
crossrefAI2026-01-16

A Radiomics-Based Machine Learning Model for Predicting Pneumonitis During Durvalumab Treatment in Locally Advanced NSCLC

Takeshi Masuda, Daisuke Kawahara, Wakako Daido, Nobuki Imano, Naoko Matsumoto, Kosuke Hamai, et al.

Introduction: Pneumonitis represents one of the clinically significant adverse events observed in patients with non-small-cell lung cancer (NSCLC) who receive durvalumab as consolidation therapy after chemoradiotherapy (CRT). Although clinical factors such as radiation dose (e.g.…

View free PDFSource page
crossrefAI2024-11-19Cited by 5

A Novel Multi-Objective Hybrid Evolutionary-Based Approach for Tuning Machine Learning Models in Short-Term Power Consumption Forecasting

Aleksei Vakhnin, Ivan Ryzhikov, Harri Niska, Mikko Kolehmainen

Accurately forecasting power consumption is crucial important for efficient energy management. Machine learning (ML) models are often employed for this purpose. However, tuning their hyperparameters is a complex and time-consuming task. The article presents a novel multi-objectiv…

View free PDFSource page
crossrefAI2024-11-14

SIBILA: Automated Machine-Learning-Based Development of Interpretable Machine-Learning Models on High-Performance Computing Platforms

Antonio Jesús Banegas-Luna, Horacio Pérez-Sánchez

As machine learning (ML) transforms industries, the need for efficient model development tools using high-performance computing (HPC) and ensuring interpretability is crucial. This paper presents SIBILA, an AutoML approach designed for HPC environments, focusing on the interpreta…

View free PDFSource page
crossrefAI2024-08-06Cited by 9

Optimizing Curriculum Vitae Concordance: A Comparative Examination of Classical Machine Learning Algorithms and Large Language Model Architectures

Mohammed Maree, Wala’a Shehada

Digital recruitment systems have revolutionized the hiring paradigm, imparting exceptional efficiencies and extending the reach for both employers and job seekers. This investigation scrutinized the efficacy of classical machine learning methodologies alongside advanced large lan…

View free PDFSource page
crossrefAI2025-07-09Cited by 2

Interactive Mitigation of Biases in Machine Learning Models for Undergraduate Student Admissions

Kelly Van Busum, Shiaofen Fang

Bias and fairness issues in artificial intelligence (AI) algorithms are major concerns, as people do not want to use software they cannot trust. Because these issues are intrinsically subjective and context-dependent, creating trustworthy software requires human input and feedbac…

View free PDFSource page