arxivcs.LGcs.RO2026-06-30
Warp RL: Reshaping Base Policy Distributions for Dynamics Adaptation
Ethan Hirschowitz, Fabio Ramos
Residual reinforcement learning adapts a pretrained robot policy by learning an additive correction to its actions. While effective when adaptation amounts to shifting the base policy's action distribution, additive corrections cannot change the distribution's shape, scale, or st…