arxivcs.RO2026-07-14
Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences
Taehyung Kim, Gwangmo Lee, Minjun Chang, Sunghyun Lim, Jongeun Choi
Aligning robot policies with human preferences is essential for deployment to diverse end users. In per-user alignment approach, preference feedback is often sparse, so learning becomes unstable and vulnerable to human preference noise, and a growing number of individualized poli…