arxivcs.ROcs.AI2026-06-30
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
Hao Sun, Yu Song, Shiyu Teng, Ziwei Niu, Yen-Wei Chen
VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suffer from intrinsic limitations: temporal myopia that discards historical dynamics, reasoning gaps bet…