arxivcs.AIcs.CL2026-07-22
Rewarding Better Thinking for LLM Preference Alignment
Xubo Liu, Wenya Guo, Ruxue Yan, Xinying Qian, Ying Zhang
LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome-level, mainly evaluating the final response whi…