arxivcs.LGcs.CL2026-07-21
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs
Seunghyun Lee, Dongyoon Han, Sangdoo Yun
Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal training); both pay a tax in adjacent-domain competence or over-refusal. We argue that the right op…