CORTEXA
← Browse
arxivcs.AI2026-07-14

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

Sen Yang, Yuen-Hei Yeung

Aligned language models routinely misreport under non-evidential pressure: they cave to a confident user, yet fail to revise when genuine evidence arrives. We cast this as a failure of internal incentive-compatibility and study the two demands, resist (ignore forbidden pressure) and update (follow licensed evidence), on a Bayesian-witness benchmark with known posteriors, where the same user disagreement is evidence or pressure purely by stated source reliability, removing the evidence/pressure confound by construction. Using interchange interventions rather than probes, we causally localize low-rank report coordinates for answer, confidence, and caveat, establishing causal sufficiency at a late intervention site rather than uniqueness or necessity, with a causal cross-talk matrix showing strong own-coordinate control and only small cross-effects (partial functional disentanglement). We then introduce a training-free counterfactual report-coordinate (CRC) clamp that references the model's own report under an incentive-neutralized counterfactual of the prompt. The two-pass full-window clamp attains resist and update of $1.00$ jointly (Wilson 95% CI $[0.99,1.00]$; the rank-16 projection alone reaches $0.88/0.90$), which we read as a causal certificate and upper bound under a constructible reference, not a claim of a deployed solution. Tested global decoding and fixed-direction steering trade one objective against the other, and resist-only training collapses updating to $0.01$. The deployable single-pass compilation is lossy ($0.73/0.97$). The mechanism and the clamp reproduce across three model families and transfer to a natural sycophancy benchmark with significant paired improvements. Our contribution is the interface and certification method: activation-level counterfactual incentive-invariance as a structural primitive for internal incentive-compatibility.

View free PDFSource page

Related papers

arxivcs.AI2026-06-25

TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation

Zhixiang Lu, Xiwei Liu, Sifan Song, Changkai Ji, Anh Nguyen, Jionglong Su, et al.

Transcatheter Aortic Valve Replacement (TAVR) planning requires meticulous multimodal reasoning. However, adapting Multimodal Large Language Models (MLLMs) to this high-stakes domain is severely impeded by diagnostic hallucinations, where generated text lacks anatomical grounding…

View free PDFSource page
arxivcs.CLcs.AI2026-07-14

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors

Winston Zeng, Ali Emami, Jinho D. Choi

What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers o…

View free PDFSource page
arxivcs.AIcs.HC2026-07-08

Learning social norms enhances compatibility in dynamic human-AI coordination

Yi Yang, Siyuan Liu, Xin Gao, Huamu Sun, Chao Liu, Qing Zhou, et al.

Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit expectations among interacting agents. As AI agents, including large language models (LLMs), become embedded in daily life, they incr…

View free PDFSource page