CORTEXA
← Browse

Anlan Zhang

1 paper indexed

arxivcs.LGstat.AP2026-06-26

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, et al.

AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains. When uncalibrated AI evaluations are used for model ranking, item scoring, or…

View free PDFSource page