arxivcs.LGcs.AIcs.CV2026-07-06
Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models
Dmitriy Poyarkov, Aleksei Staroverov, Aleksandr I. Panov
It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particular, RL-trained policies exhibit stronger out-of-distribution (OOD) behavior, where models trained on…