CORTEXA
← Browse

Linyang Li

1 paper indexed

arxivcs.LGcs.CL2026-07-21

H$^2$SD: Hybrid Hindsight Self-Distillation

Qiye Cai, Yichuan Ma, Linyang Li, Peiji Li, Yongkang Chen, Qipeng Guo, et al.

Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level guidance. Existing self-distillation methods add a privileged teacher but typically assign it a fixe…

View free PDFSource page