CORTEXA
← Browse

Anagha Radhakrishna Palandye

1 paper indexed

arxivcs.LG2026-07-03

Reward Granularity in RLVR: Comparing Process and Outcome Reward Structures for Mathematical Reasoning in Small Language Models

Anagha Radhakrishna Palandye, Rebecca Glick, Osheen Kaul

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for improving mathematical reasoning in language models. Yet most RLVR work rewards only the final answer (outcome-based rewards), leaving the impact of step-level process supervision (proce…

View free PDFSource page