CORTEXA
← Browse

Manling Li

1 paper indexed

arxivcs.ROcs.AI2026-06-28

Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering

Hao Wang, Jiuzhou Lei, Dayou Li, Bangya Liu, Minghui Zheng, Manling Li, et al.

Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a policy trained on diverse handover demonstrations may learn to pass a knife blade-first. Standard remedie…

View free PDFSource page