CORTEXA
← Browse

Anna Gavenčiak

1 paper indexed

arxivcs.AIcs.LG2026-06-28

Safety from Honesty in a Disinterested AI Predictor

Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak, Michael Cohen, Rory Svarc, Damiano Fornasiere, et al.

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate t…

View free PDFSource page