arxivcs.CV2026-07-17
How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA
Navya Gupta, Bingjie Xu, Avinash Anand, Timothy Liu, Zhengchen Zhang
Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attribute verification. Despite strong aggregate performance, the mechanistic basis of VLM failures on t…