arxivcs.CVcs.AI2026-07-16
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
Wei Li, Peijin Jia, Yuan Ma, Xuefeng Jiang, Titong Jiang, Sheng Sun, et al.
Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods pre…