arxivcs.CVcs.LG2026-07-14
Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks
Wenhao Zhang, Zhongliang Zhou, John Kang, Sheng Li
Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answering (VQA) benchmarks. We audit these claims and find them fundamentally compromised by data leakage at two hierarchical levels:…