arxivcs.LG2026-07-10
Nonparametric Bayesian Inverse Reinforcement Learning with Data-Parallel Gibbs Sampling
Sai Anirudh Katupilla, Shreeya Dasa Lakshminath
Inverse Reinforcement Learning recovers reward functions from expert demonstrations, but standard formulations assume that all demonstrations come from a single expert. When demonstrations are pooled from multiple experts with distinct preferences, parametric methods recover an a…