arxivcs.CVcs.LG2026-07-03
WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling
Zelin Zhao, Min Shi, Bo Yuan, Haotian Xue, Jialuo Li, Lama Moukheiber, et al.
World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-World (VLAW) modeling. Meanwhile, unified vision-language models have demonstrated strong multimodal gen…