Explainable XGBoost model and nomogram for risk factor identification and risk prediction in cerebral small vessel disease: a machine learning-based retrospective cohort study
Xi Zhu, X J Liu, Xujie Wang, Yuyi Fan, Ming Chen
Background Cerebral small vessel disease (CSVD) is a common, clinically significant vascular disorder that frequently leads to cognitive impairment, dementia, and poor overall prognosis. Owing to its complex hemodynamic characteristics and multifactorial pathophysiology, early identification of individuals at high risk for CSVD remains a clinical challenge. This study aimed to develop and validate an interpretable machine learning (ML) model for predicting the occurrence of CSVD. Methods We retrospectively enrolled 1,640 adult patients treated at the Fifth Affiliated Hospital of Xinjiang Medical University between September 2019 and December 2024. Twenty-three candidate variables (demographics, vitals, biomarkers, comorbidities) were evaluated. Feature selection was performed using least absolute shrinkage and selection operator (LASSO) regression, followed by stepwise backward elimination in multivariable logistic regression. Six supervised ML algorithms (DT, KNN, LR, LightGBM, XGBoost, SVM) were compared. Performance was assessed using ROC curves, calibration plots, and decision curve analysis (DCA). The optimal model was interpreted using SHapley Additive exPlanations (SHAP), and a bedside clinical nomogram was constructed. Results Ten independent predictors were identified: blood glucose, history of hypertension, systolic blood pressure, age, triglycerides, history of stroke, cystatin C, C-reactive protein, homocysteine, and body mass index. Among all models, XGBoost demonstrated the best performance, with an AUC of 0.968 in the training cohort and 0.938 in the validation cohort. Calibration plots and DCA confirmed its clinical utility. The derived nomogram demonstrated strong prognostic discrimination ( p < 0.0001). The XGBoost model achieved an accuracy of 88.0%, sensitivity of 80.9%, specificity of 93.8%, and an F1 score of 0.86, corresponding to a 5.4-percentage-point gain in AUC over logistic regression. Ten-fold cross-validation confirmed this ranking, with a mean AUC of 0.934 ± 0.016. Conclusions We validated an interpretable XGBoost-based ML model that facilitates early risk stratification and targeted interventions for CSVD. Because the model relies only on routinely collected, low-cost variables and open-source software, it is readily transferable to resource-limited settings; future work will focus on prospective, multicentre external validation and on embedding the nomogram into electronic-health-record decision support.