arxivcs.ROcs.CV2026-07-02
The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection
Jincheng Tang, Yilong Zhu, Zhengyuan Xie, Jiang-Jiang Liu, Jiaxing Zhang
Vision-Language-Action (VLA) models have shown remarkable promise in generalized robotic manipulation. However, their spatial generalization remains fragile. We argue that simply increasing the number of viewpoints is insufficient. Models often fall into the trap of Shortcut Lear…