arxivcs.CLcs.AIcs.LG2026-07-22
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent…