Comparative diagnostic accuracy of radiomics, deep learning, and hybrid AI for invasiveness stratification of pulmonary ground-glass nodules: a systematic review and meta-analysis
Ning Dong, Zhongwei Li, Zhenjie Cong, Xinru Ba, Jing Qin, Zhuxiao Lin, Xiaolin Wang, Peng Liang
Background Accurate preoperative stratification of ground-glass nodules (GGNs) into invasive versus non-invasive adenocarcinomas is essential for optimizing surgical management and preventing overtreatment. Radiomics, deep learning, and hybrid models have shown potential for noninvasive histopathologic prediction, but their comparative diagnostic performance and generalizability remain unclear. This study systematically compares the diagnostic accuracy, methodological quality, and clinical applicability of radiomics, deep learning, and hybrid AI approaches for differentiating invasive from non-invasive lung adenocarcinomas manifesting as GGNs. Methods A comprehensive search of PubMed, Embase, Cochrane Library, and Web of Science was conducted through 24 September 2025 (PROSPERO CRD42025115051). Twenty-two studies (9,768 nodules) met inclusion criteria, evaluating AI models with exact segmentation or embedded nodule localization and histopathologic confirmation. Results Hybrid models demonstrated higher pooled diagnostic performance (AUC 0.94, sensitivity 0.91 [95% CI 0.88–0.93], specificity 0.87 [0.83–0.90]) than deep learning (AUC 0.92) and radiomics (AUC 0.89). At a 20% pre-test probability, a positive hybrid AI result increased post-test probability of invasiveness to approximately 59%, while a negative result decreased it to approximately 5% (approximate pooled LR + 6.0; LR– 0.22). However, effect estimates showed substantial heterogeneity (I² > 60%) and publication bias suggested sensitivity inflation (3–5%). Most studies were single-center and retrospective, with 82% conducted in East Asia and limited external validation. Conclusion Hybrid AI showed higher pooled accuracy compared with radiomics or deep learning alone, but the certainty of evidence is low–moderate due to inconsistency, publication bias, and restricted population diversity. AI-based GGN risk stratification should currently complement, rather than replace, multidisciplinary clinical decision-making until validated in prospective multicenter trials. Systematic review registration: https://www.crd.york.ac.uk/prospero/display_record.php?RecordID=CRD42025115051 , identifier CRD42025115051.