arxivcs.LGcs.AI2026-07-18
Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators
Alireza Furutanpey, Schahram Dustdar
Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise. Surprise minimization is scoped by design to ``unstable'' environments. Prediction-error curiosi…