arxivcs.CV2026-07-23
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
Han Li, Si Liu, Zehao Huang, Dongxin Lyu, Longfei Xu, Jiahui Fu, et al.
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation of the real world, such as spatial perception and dynamic r…