arxivcs.SEcs.AI2026-07-15
Copy-on-Write Scoring: Application-Specific Agent Evaluations
Trustworthy deployment of LLM-based agents in software systems requires evaluating how they perform on application-specific workflows, with enough granularity to localize where they succeed and fail. Yet existing agent evaluation mechanisms are limited: benchmarks have low constr…