arxivcs.LG2026-07-08
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma
Chongyu Fan, Pengfei Liu, Jingjia Huang, Sijia Liu, Yi Lin
Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of large language models (LLMs). To achieve sample efficiency, modern RL frameworks rely on importance sampling (IS). However, these algorithms suffer from an exploration…