CORTEXA
← Browse

Tanush Chopra

1 paper indexed

arxivcs.AIcs.LG2026-07-09

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky, Tanush Chopra, Victoria Krakovna

Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. While effective in standard scenarios, recent work highlights that LLMs remain vulnerable to persuasio…

View free PDFSource page