arxivcs.LGcs.CLstat.ML2026-06-28
Quantifying Ranking Uncertainty in LLM Benchmarks
Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks. Rank confidence intervals were recently introduced as a method to quantify the uncertainty in these rankings by aggregating pairwise hypothesis tests. In this wor…