CORTEXA
← Browse

Sandra Wachter

1 paper indexed

arxivcs.CLcs.AI2026-06-27

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

Will Hawkins, Kaivalya Rawal, Jonathan Rystrøm, Stratis Tsirtsis, Zihao Fu, Greta Warren, et al.

Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown that this increase in capability comes with a cost: it can increase a model's tendency to respond to unsafe adversarial prompts, eve…

View free PDFSource page