arxivcs.RO2026-06-29
STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning
Zhihao Liu, Qiuyi Gu, Yitao Wang, Dongming Qiao, Yixian Zhang, Shuaihang Chen, et al.
Real-world robot learning increasingly relies on heterogeneous data, but demonstrations and rollouts often mix useful progress with stalls, corrections, and suboptimal behavior. Effective policy learning therefore requires frame-level advantages that distinguish reliable local pr…