arxivcs.LGcs.AIcs.CLcs.CV2026-07-05
DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
Silin Gao, Hao Zhao, Zeming Chen, Sepideh Mamooler, Antara Raaghavi Bhattacharya, Qiyu Wu, et al.
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multiple levels of dynamic constituents, such as actions taken in the visual sequence, and the associated…