arxivcs.LG2026-07-20
Theoretical Foundations of $\max$@$k$ Reinforcement Learning
Riccardo Poiani, Martino Bernasconi, Andrea Celli
Reinforcement Learning is a cornerstone technique for modern large reasoning models. Usually, for difficult tasks such as code generation and theorem proving, the agent is evaluated by generating $K$ responses rather than sampling a single response, and performance is then measur…