CORTEXA
← Browse

Ruikang Zhao

1 paper indexed

arxivcs.CLcs.AIcs.LG2026-06-30

SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing

Ruikang Zhao, Zhenting Wang, Han Gao, Ligong Han

Reinforcement learning for diffusion large language models (dLLMs) has largely moved to trajectory-aware methods. The current state of the art, TraceRL, holds that random masking is mismatched with the model's inference trajectory, and it reconstructs that trajectory during train…

View free PDFSource page