arxivcs.AIcs.LG2026-07-02
Evidence-State Rewards for Long-Context Reasoning
Long-context reasoning requires models to locate, revise, and synthesize evidence distributed across lengthy inputs. Existing long-context RL methods usually reward final answers or static evidence extraction, offering little feedback on how intermediate actions change the model'…