arxivcs.CRcs.AI2026-06-29
Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models
Arash Raftari, Mehrdad Mahdavi, Nathan Blackthorn, Andrew Arash Mahyari
Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is present. In this work, we study post hoc detoxification of backdoored LLMs in a practical setting w…