arxivcs.AI2026-07-24
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
Heyang Jiang, Henry Liu, Baharan Mirzasoleiman
Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful instantiations. However, GRPO relies on repeated generation of long chain-of-thought rollouts. Traini…