arxivcs.CLcs.AIcs.SE2026-07-06
ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents
Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails. A model that never calls a needed tool and a model that calls the tool but ignores the result can look similar under final task accuracy. We introduce ToolFail…