arxivcs.CLcs.AIcs.LG2026-07-16
Mask-Aware Policy Gradients for Diffusion Language Models
Haran Raajesh, Kulin Shah, Adam Klivans, Philipp Krähenbühl
Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-like…