arxivcs.AI2026-06-26
ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents
Qitai Tan, Zefang Zong, Yang Li, Peng Chen
Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but its gains saturate once the stud…