Development and external validation of machine learning models for predicting mortality in patients with lung cancer: a multicenter retrospective cohort study
Lihui Liu, Pan Zuo, Chunlan Yu, Xinhai Shen, Baoping Luo, Xicheng Zhang, Ting Cao, Yong Zhou, Kaiwen Hu, Cai Yn
Lung cancer remains the leading cause of cancer mortality worldwide. Accurate prognostic prediction can support clinical decision-making and resource allocation, yet many existing models use limited predictors and lack independent validation. We developed and externally validated machine-learning models to predict mortality in patients with lung cancer using routinely collected clinical, laboratory, and treatment-related variables. We conducted a multicenter retrospective cohort study including 2,783 patients with primary lung cancer from seven hospitals in China (January 2011 to end of data collection). Six centers formed the development cohort (n = 2,398; 70% training [n = 1,678] and 30% internal test [n = 720]), and one independent center served as external validation (n = 385). Predictors were selected using Boruta followed by LASSO. We compared six algorithms (logistic regression, random forest, support vector machine, gradient boosting machine, XGBoost, and neural network). Discrimination, calibration, and clinical utility were assessed using AUC, calibration plots/Brier score, and decision curve analysis. Model interpretation used SHAP. Seventeen predictors were retained from 51 candidates, spanning treatment (chemotherapy, intervention), tumor stage (T stage, M stage, stage group, distant metastasis), inflammatory markers (white blood cell count, neutrophil percentage, monocyte-to-lymphocyte ratio), nutrition/hematology (albumin, red blood cell count), tumor markers (CYFRA21-1, CEA), coagulation (PT, APTT), and electrolytes (sodium, potassium). Random forest performed best in the internal test set (AUC 0.861, 95% CI 0.835–0.886). In external validation, performance attenuated but remained acceptable (AUC 0.749, 95% CI 0.693–0.799; accuracy 72.7%; sensitivity 80.0%; specificity 49.4%), with reasonable calibration (Brier score 0.159) and net benefit across relevant thresholds. SHAP identified T stage, chemotherapy status, white blood cell count, albumin, and M stage as key contributors. We developed a web-based risk calculator. A multicenter machine-learning model integrating routinely available variables achieved moderate externally validated performance for mortality prediction in lung cancer. The accompanying web-based calculator may facilitate individualized risk stratification, pending further validation and potential recalibration in new settings.