arxivcs.CRcs.AI2026-06-26
ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
Shijing Hu, Liang Liu, Zhu Meng, Zhicheng Zhao
Large language models (LLMs) have increasingly moved from standalone text generation systems to agents that invoke external tools, access environments, and execute multi-step tasks. However, conventional function-calling benchmarks mainly evaluate task completion and API correctn…