arxivcs.AIcs.CVcs.LG2026-06-30
Harnessing Textual Refusal Directions for Multimodal Safety
Moreno D'Incà, Nicu Sebe, Massimiliano Mancini
To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feasible in Multimodal LLMs (MLLMs) as they require unsafe multimodal data, harder to collect than their…