arxivcs.LGcs.AI2026-07-03
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy
Zijun Xie, Yuyang You, Yongzhi Li, Enlei Gong, Zeyu Chen, Quan Chen, et al.
Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult. Existing scalable RL methods typically assign trajectory-level rewards uniformly across tok…