CORTEXA
← Browse

Federico Danieli

1 paper indexed

arxivcs.LGcs.AI2026-06-26

Normalized Rewards for Preference Optimization

Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald, Katherine Metcalf

Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their implicit reward model and decrease the likelihood of preferred responses. This results in a decreas…

View free PDFSource page