arxivcs.AIcs.LGstat.ML2026-07-10
ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
Matthew ffrench-Constant, Daniel Yang, Xinmeng Huang, Sanyam Kapoor
Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone is insufficient: models must also know when they are likely to be wrong. We present ConfidenceBench, a calibration benchmark that…