An interpretable machine learning model for chronic kidney disease identification among obese adults: a nationwide population-based study
HyeokJun Yang, Young Gyun Seo, Nayoung Han
Chronic kidney disease (CKD) is a major public health concern, particularly among individuals with obesity; however, population-level identification of CKD remains challenging. This study aimed to develop an interpretable machine learning model for CKD identification and to investigate associated contributing factors using nationwide data. We analyzed 37,241 obese adults aged 19–79 years from the Korea National Health and Nutrition Examination Survey (2007–2023). Feature selection was performed using recursive feature elimination with cross-validation (RFECV), and an eXtreme Gradient Boosting (XGBoost) model was developed. Model performance was evaluated using cross-validation, calibration, and decision curve analysis, and interpretability was assessed using Shapley Additive exPlanations (SHAP) and individual conditional expectation analyses. Fifteen variables were selected, among which age, metabolic comorbidities, and sex were the most influential factors associated with CKD status. The final model demonstrated moderate discriminative performance (AUROC 0.752), and the model showed acceptable calibration performance. Lifestyle-related variables, including frequency of eating out, and smoking duration showed modest contributions but were retained as potentially modifiable predictors. This interpretable machine learning approach highlights the importance of metabolic comorbidities and age in CKD identification among obese adults while providing a foundation for future longitudinal risk stratification studies.