arxivcs.SEcs.LG2026-07-08
DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
Wenqi Huang, Charley Lee, Leonard Tng, Serena Ge
DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents. Most public agentic coding benchmarks follow SWE-bench in mining merged fixes from public GitHub repositories, which creates two problems: the fixes and their discussion…