CORTEXA
← Browse

Qiang Zhu

3 papers indexed

arxivcs.AI2026-07-15

LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning

Qiang Zhu, Jiajun Wu, Longyi Wang

Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and harmful intermediate interactions. We propose LOTAPO , a self-generated process-supervision method based on backward leave-one-turn…

View free PDFSource page
arxivcs.AI2026-07-06

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training

Qiuyi Qi, Tian Liang, Mutian Bao, Jinjian Zhang, Dongnan Liu, Wei Zhou, et al.

Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to trajectory neglect, in which agents lose focus on the task goal and interaction history at intermediate ste…

View free PDFSource page
arxivcs.AI2026-07-06

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

Qiuyi Qi, Jinjian Zhang, Mutian Bao, Tian Liang, Guocong Li, Dongnan Liu, et al.

Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficiency arises from a lack of systematic mechanisms t…

View free PDFSource page