An explainable machine learning framework for academic stress classification among university students: a comparative multi-dataset study with ablation analysis and statistical significance testing
Assel Omarbekova, Aizhan Nazyrova, Gulmira Bekmanova, Zhanar Lamasheva, Albina Gibadullina, Assem Abdikarim
Academic stress among university students represents a major public health concern due to its detrimental effects on cognitive functioning, psychological well-being, and long-term academic performance. Although machine learning has increasingly been applied to stress prediction, existing studies commonly rely on single datasets, insufficiently address class imbalance, and provide limited model interpretability. This study proposes a comparative machine learning framework that integrates explainable artificial intelligence (XAI), ablation analysis, and statistical significance testing to improve the classification of student stress. Two complementary datasets were analysed. Dataset 1 included 777 university students aged 18–22 years with three stress-type categories (Eustress, Distress, and No Stress), while Dataset 2 comprised 1,100 records with 20 psychological, physiological, environmental, and academic features labelled according to three stress severity levels (Low, Medium, and High). Five supervised learning algorithms—Logistic Regression, Random Forest, Gradient Boosting, Support Vector Machine with a radial basis function kernel, and Multilayer Perceptron—were evaluated using 10-times repeated stratified 5-fold cross-validation. Severe class imbalance in Dataset 1 (91.4% Eustress) was addressed using the Synthetic Minority Over-sampling Technique (SMOTE), and its impact was quantified through a structured ablation study. Model performance differences were assessed using the Friedman test with post-hoc Nemenyi analysis, and SHAP was employed to provide global and class-specific model interpretability. The Friedman test demonstrated statistically significant differences among classifiers ( χ ² = 68.97, * p * < 0.001), with post-hoc analysis identifying Random Forest and Logistic Regression as the statistically superior-performing models. Random Forest achieved the highest overall classification accuracy (0.8909) and the lowest log loss (0.2057), whereas Logistic Regression produced the highest discriminative performance (AUC-ROC = 0.9853). The ablation study confirmed that SMOTE substantially improved minority-class prediction performance. SHAP analysis identified blood pressure, sleep quality, teacher–student relationship, and academic performance as the most influential predictors of overall stress severity, while self-esteem, sleep quality, bullying, and anxiety level were the strongest contributors to the High Stress class. The proposed framework provides a reproducible, interpretable, and statistically validated approach for multi-class stress classification among university students. By combining class imbalance correction, rigorous comparative evaluation, explainable artificial intelligence, and statistical validation, the framework addresses important methodological limitations of previous studies and offers a robust foundation for the development of early stress detection and intervention systems in higher education.