arxivcs.LGcs.AI2026-07-15
RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
Logan Mondal Bhamidipaty, Mykel Kochenderfer, Subramanian Ramamoorthy
World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset. However, they are vulnerable to model exploitation where data coverage is thin. Prior work addresses this either by collecting more exp…