arxivcs.CRcs.AI2026-07-08
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang, Longwei Wang
Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how ad…