CORTEXA
← Browse

Johannes Treutlein

1 paper indexed

arxivcs.LGcs.AIcs.CR2026-07-15

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

Jan Betley, Johannes Treutlein, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, et al.

People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluation…

View free PDFSource page