arxivcs.AI2026-07-24
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
Jiaqi Shao, Hanck Chen, Wei Zhang, Maxm Pan, Bing Luo
Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and…