arxivcs.CVcs.AI2026-06-27
X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving
Bohao Zhao, Chengrui Wei, Guangfeng Jiang, Ruixin Liu, Xuejie Lv, Liu Liang, et al.
Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive perception-action mapping. While integrating Predictive World Models (PWMs) addresses this gap, existing ap…