arxivcs.ROcs.AI2026-06-30
Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models
Lang Cao, Renhong Chen, Luyi Li, Peng Wang, Mofan Peng, Yitong Li
Vision-Language-Action (VLA) models offer a promising framework for robotic manipulation by connecting language instructions, visual observations, and continuous control. However, most existing policies remain limited by behavior cloning or supervised fine-tuning (SFT) from fixed…