CORTEXA
← Browse
arxivcs.CYcs.AI2026-06-27Cited by 0

Defeat Devices in AI Systems

Emilio Ferrara

AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts. Alignment faking, sandbagging, benchmark gaming, deceptive scheming, specification gaming, and trojans have each been documented separately, with each line of work characterizing one facet of what we argue is a single structural mechanism. We propose that this common mechanism is a defeat device, an engineering and regulatory concept long established in vehicle-emissions law and brought to broad public attention by the 2015 Volkswagen emissions case. A defeat device in an AI system has three necessary elements: a discriminator that detects evaluation context, a concealed swap that conditions behavior on detection, and a gap between eval-distribution and deployment-distribution performance on the stated evaluation criterion. We formalize this triadic test as a behavioral definition, organize documented cases along three taxonomic axes (origin, trigger, swap mechanism), propose Trigger-Axis-Aware Differential Probing (TADP) as a forensic detection protocol, and advance the claim that defeat devices can naturally emerge in current frontier AI systems without any operator engineering. We characterize naturally-emerging defeat devices as potentially one of the harmful emerging phenomena that AI safety practice should monitor and test for systematically. Implications for evaluation methodology, post-training pipeline design, interpretability research priorities, and AI governance follow.

View free PDFSource page

Related papers

arxivcs.HCcs.AIcs.CYcs.ETcs.RO2026-07-14

Practical Judgment, Virtue, and Intuition in the Use of Opaque AI-Enabled Systems

Nathan G. Wood, Andrew P. Rebera

AI-enabled systems are seeing increasing deployment across numerous domains, with many being "black boxes" with respect to core functions and capabilities. I.e., many systems take inputs and give outputs, but without users having any ability to see how the former lead to the latt…

View free PDFSource page
arxivcs.HCcs.AIcs.CYcs.ETeess.SY2026-07-01

AI, Trust, and Teaming: The Humans-as-Handlers Approach for Autonomous and Opaque AI Systems

Nathan G. Wood

Artificial intelligence (AI) is becoming ubiquitous, and across domains, increasingly autonomous systems are carrying out tasks which raise significant ethical and legal challenges which demonstrate a need for strong human-machine teams rooted in trust. In this article, I argue t…

View free PDFSource page
arxivcs.CYcs.AI2026-06-30

A Technical Typology of AI Systems in Public Administration

Jonathan Rystrøm, Chris Schmitz, Nathan Davies, Gerhard Hammerschmid, Albert Meijer, Chris Russell

Research on artificial intelligence (AI) in the public sector often treats "AI" as a single category, neglecting technical distinctions between different AI systems. But these distinctions affect how different systems impact core public values like accountability, procedural just…

View free PDFSource page
arxivcs.CYcs.AIcs.HC2026-07-21

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Gjergji Kasneci, Enkelejda Kasneci

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rathe…

View free PDFSource page
arxivcs.AIcs.CYcs.HC2026-07-20

The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems

Samuel Presgraves

Existing AI measurement frameworks quantify cognitive capability, task automation, or catastrophic risk, but none measure autonomous agency: the extent to which a system behaves in a self-directed way. A system can saturate capability benchmarks while remaining entirely reactive,…

View free PDFSource page
arxivcs.CYcs.AIcs.CLcs.LG2026-07-18

A Method for Learning Value Systems in Generative AI

Andrés Holgado-Sánchez, Holger Billhardt, Sascha Ossowski

Value-aware AI systems require explicit computational representations of human values (groundings) and their aggregation into value systems in order to align their decisions with ours. As such representations are difficult to elicit, value learning seeks to infer them by observin…

View free PDFSource page