arxivcs.CRcs.AI2026-07-02
Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring
William Hackett, Peter Garraghan
As Large Language Models (LLMs) and agentic systems become integrated into real-world applications, ensuring their safety and security is critical. Guardrail systems that detect and block malicious instructions sent to and from an LLM are an essential component of AI security. Ho…