CORTEXA
← Browse

Yuan chen

2 papers indexed

arxivcs.LGcs.CVeess.SY2026-07-20

JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models

Ruiyi Ding, Jie Li, He Kang, Ziyan Liu, Chengru Song, Yuan chen

Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences. While successful in large language models~\cite{shao2024deepseekmathpushinglimitsmathematical}, its extension to diffusion and flow matc…

View free PDFSource page
arxivcs.CLcs.LG2026-07-08

R^3: Advertisement Compliance Rectification via Group-Relative Experience Extractor and Curriculum Reinforcement

Yuan Chen, Zhenyu Hu, Mengge Xue, Te Cao, Liqun Liu, Peng Shu, et al.

Rigorous content moderation is crucial for online advertising but leads to millions of daily rejections. This scale renders manual rectification infeasible, particularly for video advertisements. However, existing safety-driven methods often suffer from aggressive over-editing, w…

View free PDFSource page