A benchmark study of vision and pathology foundation models for computational pathology
Rohan Bareja, Francisco Carrillo‐Pérez, Yuanning Zheng, Marija Pizurica, Tarak Nath Nandi, Lu Tian, Jeanne Shen, Ravi Madduri, Olivier Gevaert
To advance precision medicine in pathology, artificial intelligence (AI)-driven foundation models must generalize across diverse datasets, tissues, and clinical tasks. However, their comparative performance and generalizability in computational pathology remain incompletely characterized. Here, we benchmark 32 AI foundation models across four categories, including general vision models (VM), general vision-language models (VLM), pathology-specific vision models (Path-VM), and pathology-specific vision-language models (Path-VLM), using slide- and patch-level tasks from The Cancer Genome Atlas (TCGA), Clinical Proteomic Tumor Analysis Consortium (CPTAC), external benchmarking datasets, and out-of-domain datasets. Across TCGA tasks, Path-VMs consistently rank among the strongest performers. Evaluation across CPTAC and out-of-domain datasets reveals more nuanced generalization behavior, with model rankings showing modest but consistent shifts across datasets and task categories. Pairwise statistical comparisons indicate that differences among top-performing models are often small and task dependent. Path-VMs outperform Path-VLMs and remain competitive with VMs. Model size and pretraining dataset scale do not consistently predict downstream performance. Finally, late decision-level ensembling improves aggregate performance across external datasets and tissue types, highlighting complementary strengths across foundation models. PathBench: https://pathbench.stanford.edu/ The comparative performance and generalisability of pathology foundation models remain largely unexamined. Here, the authors benchmark 32 AI pathology foundation models across large cancer datasets, showing that generalisation in computational pathology is heterogeneous and task-dependent regardless of dataset scale, but ensemble-based approaches can combine the strengths of different models.