arxivcs.CRcs.AIcs.HC2026-07-20
Towards an Automated Test of LLM Security Knowledge
Shufan Chai, Liangliang Sun, Jessica Staddon
Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on identifying areas in which LLM security ``k…