CORTEXA
← Browse

Bowei He

2 papers indexed

arxivcs.LGcs.CL2026-07-15

Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning

Bowei He, Yankai Chen, Xiaokun Zhang, Xue Liu

Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes. State-of-the-art algorithms such as PPO, RLOO, and GRPO inherit their rollout topology from RLHF: for each prompt, N independent tra…

View free PDFSource page
arxivcs.LGcs.AIcs.CL2026-07-15

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

Ye Yuan, Weien Li, Rui Song, Zeyu Li, Haochen Liu, Xiangyu Kong, et al.

Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed,…

View free PDFSource page