CORTEXA
← Browse
crossrefAI2026-07-18Cited by 0

Predicting Student Stress Using Machine Learning Ensemble Models: A Multi-Criteria Comparison with Explainable Artificial Intelligence Analysis

Daniel Cristóbal Andrade-Girón, William Joel Marin-Rodriguez, Marcelo Gumercindo Zuñiga-Rojas, Abrahan Cesar Neri-Ayala, Edgar Tito Susanibar-Ramírez, Miguel Angel Aguilar-Luna-Victoria

Student stress is a significant mental health issue in educational settings; therefore, developing reliable, calibrated, and interpretable predictive models can support the classification of observed stress levels. This study analyzed the public Student Stress Factors dataset, comprising 1100 records, 20 predictors, and one target variable, using a supervised machine learning pipeline designed to reduce information leakage. The pipeline included stratified data partitioning, encapsulated preprocessing, nested cross-validation restricted to the training data, and independent holdout evaluation. Nine ensemble and boosting algorithms for tabular data were compared: AdaBoost, Gradient Boosting, Random Forest, Extra Trees, Bagging, Voting, Stacking, XGBoost, and LightGBM. Model performance was assessed using key discrimination and calibration metrics, together with the nonparametric Friedman test for statistical comparison. Gradient Boosting achieved the best average performance in nested cross-validation, with an accuracy of 89.55 ± 3.16%, F1-weighted of 89.54 ± 3.17%, MCC of 0.845 ± 0.047, and ROC-AUC weighted of 98.59 ± 0.92%. XGBoost and LightGBM showed comparable performance. In the independent holdout set, the final calibrated model maintained robust predictive performance, achieving an accuracy of 0.8818, F1-weighted of 0.8818, MCC of 0.8237, and ROC-AUC weighted of 0.9861. Although the overall results indicate stable and high predictive performance, the Friedman test did not identify statistically significant differences among the algorithms, χ2 = 10.953, p = 0.204. Therefore, model selection should consider not only predictive accuracy but also computational efficiency, calibration, interpretability, and implementation feasibility. Despite the internal stability of the pipeline and satisfactory holdout performance, the public and cross-sectional nature of the dataset limits causal inference and model transferability. Consequently, external and prospective validation is required before integration into institutional early warning systems.

View free PDFSource page

Related papers

crossrefAI2026-07-24

A Compact Deep Learning Framework for Potato Leaf Disease Classification Across Controlled/Uncontrolled Environments

Omneya Attallah

Potato leaf disease poses a significant threat to global food security, causing substantial crop losses that jeopardise agricultural productivity and farmers’ livelihoods worldwide. Existing automated detection frameworks suffer from several persistent limitations, including over…

View free PDFSource page
openalexAI2026-07-23

Let’s Code with GenAI: Exploring K-12 Teachers’ Self-Efficacy, Value Beliefs, and Coding Performance

Suk Tae Kang, Wanju Huang

While computational thinking (CT) is increasingly vital in K-12 education, teaching it through text-based coding remains challenging for teachers. To address this gap, this study presents and evaluates a self-paced professional development (PD) module, “Let’s Code with GenAI,” cr…

View free PDFSource page
crossrefAI2026-07-12

eGFR-AI: A Stacked Machine-Learning Model for Early Postoperative Kidney Function Prediction—A Pilot Study

Eva Brenner, Luka Bulić, Vilena Vrbanović Mijatović

Background: Postoperative kidney dysfunction is a common and serious complication in surgical patients. Kidney function is typically assessed using the estimated glomerular filtration rate (eGFR), most often calculated with the CKD-EPI equation based on serum creatinine. While se…

View free PDFSource page
crossrefAI2026-07-01

Towards Data-Driven Weather Intelligence in Palestine: A Multi-Station Benchmark of Classical Machine Learning and Deep Learning Models

Mohammad Odeh, Ahmad Hasasneh

Precise weather forecasting plays a critical role in sectors such as agriculture, transport, energy management, and climate change adaptation, and machine learning and deep learning algorithms have been widely used for data-driven time series forecasting problems. In this work, w…

View free PDFSource page
crossrefAI2026-06-01

Beyond Vital Signs: A Machine Learning Model Using Comprehensive Triage-Time Data to Detect Undertriage in Emergency Department Patients

Kyungman Cha, Sohee Lee, Jaekwang Shin, Jee Yong Lim

Undertriage—the misclassification of acutely ill patients into low-acuity triage categories—is a persistent patient safety concern, and prior machine learning approaches restricted to vital signs have yielded modest predictive performance. We hypothesized that this ceiling reflec…

View free PDFSource page
crossrefAI2026-05-27

Machine Learning Approaches for Filtering Organometallic Reactions: A Comparative Study of Molecular Descriptors

Walter Bonke Mahlangu, Taurai Hungwe, Nomasonto Rapulenyane, Somandla Ncube

Organometallic chemistry deals with the synthesis, structure, reactivity, and applications of compounds containing metal–carbon covalent bonds. In recent years, there has been a growing interest in predicting the catalytic activity of organometallics using machine learning. Howev…

View free PDFSource page