CORTEXA
← Browse

Niels Warncke

2 papers indexed

arxivcs.LGcs.AIcs.CR2026-07-15

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

Jan Betley, Johannes Treutlein, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, et al.

People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluation…

View free PDFSource page
arxivcs.AI2026-06-29

Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors

Maxime Riché, Daniel Tan, Vili Kohonen, Niels Warncke

Inoculation prompting is a selective-generalization technique used against Emergent Misalignment. We introduce inoculation adapters (IA), a family of methods that similarly reduce the optimization pressure to learn undesired traits by strengthening those traits during training. I…

View free PDFSource page