arxivcs.CLcs.CVcs.IRcs.MM2026-06-30
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
Qian Ma, S M Rayeed, Charles V. Stewart, Qiong Wu, Yao Ma
Knowledge-Based Visual Question Answering (KB-VQA) aims to evaluate whether Visual Language Models (VLMs) can retrieve, ground, and reason over external structured knowledge beyond visual evidence. In practice, answer accuracy is widely adopted as the primary evaluation metric, i…