arxivcs.AIcs.CVcs.RO2026-07-16
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models
Yufeng Ji, Wenhao Tang, Haoyi Niu, Koushil Sreenath, Yi Wu, Zhongyu Li
Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force that shapes inherited multimodal representations. We show that this shaping has a dual effect: it is…