arxivcs.LGcs.AI2026-06-25
GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning
Ting Zhou, Zhenqing Ling, Yiyang Zhao, Ying Shen, Daoyuan Chen
Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency: within a batch, a small set of high-reward rollo…