arxivcs.LGcs.CVeess.SY2026-07-20
JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models
Ruiyi Ding, Jie Li, He Kang, Ziyan Liu, Chengru Song, Yuan chen
Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences. While successful in large language models~\cite{shao2024deepseekmathpushinglimitsmathematical}, its extension to diffusion and flow matc…