CORTEXA
← Browse

Ofer Mendelevitch

1 paper indexed

arxivcs.AI2026-07-23

GuardianAgentBench: Where Agents Fail and How to Guard Them

Vishal Ishwar Naik, Chenyu Xu, Donna Dong, Hussein Hassan, Abhishek Pradhan, Ofer Mendelevitch, et al.

As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical. We present GuardianAgentBench (GABench), a benchmark of 580 scenarios across six domains evaluated on three…

View free PDFSource page