arxivcs.CLcs.AI2026-07-05
dOPSD: On-Policy Self-Distillation for Diffusion Language Models
Phuong Tuan Dat, Qi Li, Xinchao Wang
Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers fro…