Machine learning for chemotherapy decision-making in breast cancer using large language model
Md Serajun Nabi, Dema Yuden, Thinley Yeshey Choden, S. M. Asiful Islam Saky, Hasanul Bannah, Md Sabbir Hossen, Mohammad Faizal Ahmad Fauzi, Hezerul Bin Abdul Karim
Introduction Breast cancer chemotherapy decision-making remains challenging due to biological heterogeneity and variability in clinical practice. This study proposes a hybrid framework integrating machine learning (ML), causal reasoning, and large language models (LLMs) to improve treatment recommendations. Methods Using the METABRIC dataset, eleven pre-treatment clinicopathologic variables were selected. A Random Forest classifier was developed and compared with baseline ML models. Individualized treatment benefit was estimated through inverse probability-weighted causal survival analysis, while GPT-4 was employed using few-shot prompting to generate clinical rationales. Results The Random Forest achieved an AUC of 0.91, outperforming benchmark models. Causal analysis identified heterogeneous treatment benefits and patient groups where chemotherapy could potentially be deprioritized. GPT-4 showed moderate agreement with the Random Forest (Cohen's κ = 0.13) while consistently highlighting clinically relevant factors. Uplift-based ML policies outperformed treat-all and treat-none strategies, and GPT-4 improved interpretability through rationale-driven explanations. Discussion By combining predictive ML, causal survival modeling, and LLM-based rationale generation, the proposed framework provides a promising approach for personalized and transparent chemotherapy decision support in oncology.