arxivcs.AI2026-07-23
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization
Siwei Chen, Siqi Chen, Xupeng Miao, Bin Cui
Recent large reasoning models often develop long chain-of-thought responses during reinforcement learning (RL), resulting in high inference latency and deployment cost. Existing methods for response length control typically rely on explicit length penalties or additional control…