arxivcs.LGcs.CL2026-06-30
One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context
Yingzi Ma, Zichen Zhu, Ming Jiang, Chaowei Xiao
On-policy self-distillation (OPSD) teaches large language models new skills through a teacher that shares the student's backbone and supervises its own rollouts. Existing teachers either inject privileged context at the input -- inducing post-hoc rationalization -- or fine-tune w…