CORTEXA
← Browse

Yujun Feng

1 paper indexed

arxivcs.AI2026-07-07

Reliability-Aware LLM Alignment from Inconsistent Human Feedback

Jingyi Huang, Ruohan Zong, Yujun Feng, Liran Ma, Lanyu Shang, Yang Zhang

Reinforcement Learning from Human Feedback (RLHF) is critical for aligning Large Language Models (LLMs) with human preferences. However, its efficacy is often compromised by the inherent inconsistency and subjectivity of human annotations. Existing preference optimization framewo…

View free PDFSource page