arxivcs.AI2026-07-16
Can We Trust Item Response Theory for AI Evaluation?
Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of…