arxivcs.CV2026-07-05
Beyond Scene Priors: Fine-Grained Traffic Scene Reasoning with Benchmarking and Query-Guided Small-Object Focus
Waikit Xiu, Qiang Lu, Zian Wang, Xinjie Yang, Zhiwei Chen, Chen Sun, et al.
In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language Models (MLLMs) tend to over-attend to backgrounds, overwhelming crucial small objects during visual-language alignment, a failure…