CORTEXA
← Browse

Kristina Zhang

1 paper indexed

arxivcs.LGcs.AI2026-07-02

Out-of-Distribution Generalization of Risk Aversion in Language Models

Kristina Zhang, Junior Chinomso Okoroafor, Benjamin Maltbie, Andrew Lin, Abhitej Bokka, Elliott Thornley

Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AIs would tend to prefer low-risk, low-reward strategies like cooperation over high-risk, high-reward strategies like rebellion, limiting the d…

View free PDFSource page