Artificial Intelligence for Breast MRI Lesion Classification: A Targeted Evidence Synthesis and Meta-Analysis of Discriminative Performance and Heterogeneity
Romuald Ferré, Thad Benefield, Cherie M. Kuzmiak
Background/Objectives: The paper aimed to synthesize the diagnostic performance of artificial intelligence (AI) methods for classifying breast lesions on contrast-enhanced breast MRI and to estimate a pooled area under the receiver operating characteristic curve (AUC). Methods: This targeted evidence synthesis and meta-analysis was informed by PRISMA 2020 reporting principles where applicable. Eligible studies were drawn from an investigator-supplied corpus of 12 primary manuscripts and assessed against predefined criteria. MEDLINE/PubMed, Embase, and Web of Science were consulted through October 2025 to contextualize the literature and verify bibliographic and study details; additional database records were not screened for eligibility. We included studies applying machine learning or deep learning to contrast-enhanced breast MRI for benign-versus-malignant lesion classification and reporting an AUC on an independent test set, external validation set, or patient-wise cross-validation. AUCs were pooled on the logit scale using an inverse-variance DerSimonian–Laird random-effects model, and heterogeneity was quantified using I2. Results: Nine studies met the criteria for quantitative synthesis (evaluation-set sizes, 60–3936). The pooled random-effects AUC was 0.898 (95% CI, 0.875–0.918), with substantial heterogeneity (I2 = 88.1%) and a 95% prediction interval of 0.824–0.943, indicating that performance may vary meaningfully across settings. Conclusions: AI models showed promising discriminative performance within this targeted corpus, but substantial heterogeneity, differences in unit of analysis, approximated variance estimates, and limited external institutional validation temper confidence in generalizability. The pooled AUC should be interpreted descriptively, and future studies should prioritize rigorous multi-institutional external validation, transparent reporting, and prospective reader- or workflow-impact evaluation before routine clinical deployment.