arxivcs.SEcs.AI2026-06-26
Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
Yanuo Ma, Ben Kereopa-Yorke, Ben Schultz
Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity problems, and a passing score may not show whether the requested task was delivered. We study both problems. In a controlled code-as-spe…