Interpretable machine learning models for pediatric type 1 diabetes risk assessment
Yuda Syahidin, Nur Ulfa Maulidevi, Cissy Rachiana Sudjana Prawira, Kridanto Surendro
This study identifies clinically meaningful predictors of pediatric type 1 diabetes (T1D) risk from fully anonymized retrospective health records to support early-risk screening and clinician-facing decision support. The dataset comprises anthropometric, biochemical, autoimmune, and family-history variables routinely available in pediatric care. We propose a hybrid voting–based feature evaluation framework that aggregates evidence from multiple feature selection runs by combining selection frequency, rank-weighted importance, and AUC-weighted performance, yielding stable and interpretable feature scores. The resulting rankings highlight BMI, pancreatic markers, IA-2A, ZnT8A, and paternal diabetes history as leading predictors, consistent with established mechanisms of T1D susceptibility. Models trained on the selected features demonstrate strong discrimination on held-out testing, with Gradient Boosting and Random Forest achieving accuracies of 94.13% and 93.58%, respectively, alongside consistently high AUC. By emphasizing transparent feature-level reasoning rather than diagnostic automation, the proposed framework advances explainable and trustworthy medical AI and offers a practical pathway for integration into pediatric T1D early-screening workflows.