arxivcs.AIcs.LG2026-07-14
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
Ilias Kazantzidis, Timothy J. Norman, Yali Du, Christopher T. Freeman
We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown and no suitable reward function is available. In the context of safety-critical environments, we consider traditional reinforceme…