arxivcs.LGcs.AI2026-07-21
REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
Yunjie Chen, Xiaoxin Chen, Fang Wang
Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models (LLMs). However, continuing to scale it across vast task domains of interest remains challenging in…