arxivstat.MEcs.AIecon.EM2026-06-29
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data
Xinrui Ruan, Zhenyu Zhao, Waverly Wei, Yueshan Zhang, Zeyu Zheng, Sui Huang, et al.
Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Organizations thus often turn to collecting vast but noisy "silver" labels from crowdsourced workers or…