arxivcs.SEcs.AIcs.LGcs.PF2026-07-16
AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows
Tejas Singh Anand, Yuet Ying Christina Wang, Wanting Jiang, Steve Masson, Tian Zheng, Bingjie Zhou
Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a domain task. As skill repositories grow, developers need automated quality signals on every change, yet evaluation today is largely anecdotal…