arxivcs.CVcs.MM2026-07-11
What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks
Guanhua Ye, Niu Jingbin, Yan Li, Meiyu Liang, Zhe Xue, Yingxia Shao, et al.
Short-answer VQA benchmarks conflate two distinct quantities: whether a model's answer is semantically correct, and whether that answer matches the surface form expected by the automatic evaluator. We study this conflation across six vision--language models and six benchmarks, us…