arxivcs.LGcs.AI2026-07-12
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study this pass@k inversion: after training, the policy may solve fewer distinct problems than its base model at large $k$. The failure con…