arxivcs.ROcs.AI2026-07-10
More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning
Yi Li, Alexandre Chapin, Liming Chen, Jan Peters, Alap Kshirsagar
Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid. Both mix task-relevant and task-irrelevant features. Object-centric slot representations are a structured alternative: they group features into a few p…