CORTEXA
← Browse
crossrefMachine Learning and Knowledge Extraction2026-06-12Cited by 0

Do Foundation Models Truly Outperform Domain-Specific Models? Evidence from Digital Pathology

Chaima Ben Rabah, Ahmed Serag

Foundation models (FMs) are increasingly proposed as general-purpose solutions for computational pathology, with the potential to simplify clinical artificial intelligence deployment by reducing the need for task-specific architectures. However, their reliability across cancer domains with distinct morphological characteristics remains unclear, limiting confidence in real-world clinical use. We benchmarked seven general-purpose pathology FMs and three domain-specific FMs across eleven patch-level datasets spanning three clinically relevant domains: pediatric hematology, prostate cancer, and breast cancer, using both linear probing and last-layer fine-tuning adaptation strategies. By jointly evaluating pediatric leukemia, male-predominant prostate cancer, and female-predominant breast cancer, this study is, to our knowledge, the first to explicitly examine specialist-versus-generalist FM behavior across age- and sex-stratified cancer populations. Performance differences were strongly domain dependent. In hematology, the specialist FM DINOBloom matched and, in several datasets, marginally exceeded leading generalist models (AUC 0.990–0.999 vs. GigaPath 0.981–1.000), suggesting advantages for highly distinctive cellular morphology. In prostate cancer grading, the generalist FM UNI2-h consistently outperformed the specialist HistoEncoder (AUC 0.956–0.977 vs. 0.908–0.964). In breast cancer, UNI2-h achieved the best overall performance across all tasks. No publicly available breast-cancer-specific FM currently exists for direct comparison; therefore, breast cancer results characterize general FM transferability rather than specialist-versus-generalist differences. Importantly, cross-dataset experiments revealed substantial performance degradation under dataset shift in both prostate and breast cancer, indicating that current FMs are not yet robust enough for heterogeneous multi-site clinical use. These findings support the use of generalist FMs as efficient backbones for well-characterized single-site, patch-level tasks, while challenging the assumption that high benchmark performance necessarily reflects true clinical readiness and demonstrating that pathology FMs are not uniformly superior to specialist models.

View free PDFSource page

Related papers

crossrefMachine Learning and Knowledge Extraction2025-02-10Cited by 28

Investigating the Performance of Retrieval-Augmented Generation and Domain-Specific Fine-Tuning for the Development of AI-Driven Knowledge-Based Systems

Róbert Lakatos, Péter Pollner, András Hajdu, Tamás Joó

Generative large language models (LLMs) have revolutionized the development of knowledge-based systems, enabling new possibilities in applications like ChatGPT, Bing, and Gemini. Two key strategies for domain adaptation in these systems are Domain-Specific Fine-Tuning (DFT) and R…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-11-25Cited by 6

A Four-Dimensional Analysis of Explainable AI in Energy Forecasting: A Domain-Specific Systematic Review

Vahid Arabzadeh, Raphael Frank

Despite the growing use of Explainable Artificial Intelligence (XAI) in energy time-series forecasting, a systematic evaluation of explanation quality remains limited. This systematic review analyzes 50 peer-reviewed studies (2020–2025) applying XAI to load, price, or renewable g…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-07-04

Evaluation Rigor from Graph Neural Networks to Graph Foundation Models: A Systematic Review and a Four-Axis Reporting Standard

Sergei O. Kurashkin, Vadim S. Tynchenko, Aleksei S. Borodulin, Ahmad Hammoud, Connie Tee

Graph machine learning reports steady progress across node, graph, and link prediction, across temporal and hypergraph frontiers, and across the emerging class of graph foundation models. This review asks a prior question: when a method is reported to outperform the alternatives,…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-10-23

Interleaved Fusion Learning for Trustworthy AI: Improving Cross-Dataset Performance in Cervical Cancer Analysis

Carlos Martínez, Laura Busto, Olivia Zulaica, César Veiga

This study introduces a novel Interleaved Fusion Learning (IFL) methodology leveraging transfer learning to generate a family of models optimized for specific datasets while maintaining superior generalization performance across others. The approach is demonstrated in cervical ca…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-01-12Cited by 6

Quantifying Interdisciplinarity in Scientific Articles Using Deep Learning Toward a TRIZ-Based Framework for Cross-Disciplinary Innovation

Nicolas Douard, Ahmed Samet, George Giakos, Denis Cavallucci

Interdisciplinary research (IDR) is essential for addressing complex global challenges that surpass the capabilities of any single discipline. However, measuring interdisciplinarity remains challenging due to conceptual ambiguities and inconsistent methodologies. To overcome thes…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-06-04

A Sovereign Conversational Assistant Powered by ALIA and Mistral for the AI Act Age: Architecture, Governance, and Evaluation

Alejandro Carmona-Martínez, Antonio J. Jara, Alicia Asín

Digital Twins and Living Labs are increasingly used to support conservation, safety, accessibility, and visitor experience in cultural-heritage sites. Their practical value, however, depends on interfaces that can explain heterogeneous evidence, expose provenance, and operate und…

View free PDFSource page