arxivcs.CV2026-07-11
REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation
Li Guo, Anas M. Tahir, Z. Jane Wang
Automated chest X-ray report generation has recently benefited from reinforcement learning (RL) and large language models. However, RL training often suffers from instability or limited exploration due to fixed Kullback-Leibler (KL) regularization and a static reference policy th…