CORTEXA
← Browse
arxivcs.LGcs.AI2026-07-22

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

Kavin Aravindan, Arihant Rastogi, Krishak Aneja, Aadi Prasad, Saiyam Jain, Vaishnavi Shivkumar, Ponnurangam Kumaraguru

Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors may induce over-refusal on benign prompts. We introduce OPIUM (Optimizing Protected Injections via Utility Manifolds), a training-free method for sanitizing steering vectors through representation matching. Given reference behaviors on two prompt sets, OPIUM optimizes a new steering vector that preserves the downstream representations induced by the desired intervention while matching a safer reference behavior on prompts where the original vector fails. Across steering-externality and over-refusal settings, OPIUM improves the safety--utility tradeoff relative to vanilla steering and directional ablation, suggesting that harmful side effects of activation steering can often be mitigated directly in activation space.

View free PDFSource page

Related papers

arxivcs.ROcs.AIcs.LGeess.SYmath.OC2026-07-16

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

Jihoon Hong, Julian Skifstad, Qiyue Dai, Alice Chan, Glen Chou

World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activation…

View free PDFSource page
arxivcs.LGcs.AI2026-07-12

Conditional Optimal Bridge for Riemannian Activation Steering

Seyed Arshan Dalili, Ajay Narayanan Sridhar, Vijaykrishnan Narayanan, Mehrdad Mahdavi

Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time. While many existing methods implicitly optimize a log-density-ratio objective between desired and undesired activation distributions, they do so heuristica…

View free PDFSource page
arxivcs.LGcs.AImath.OC2026-07-07

Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization

Sounaq Das, Tanmay Sen, Raghu Nandan Sengupta, Aditya Gupta

Portfolio optimization under uncertainty is inherently a multi-objective decision problem involving complex interactions among return, risk, market dynamics, and practical investment constraints. Existing reliability based portfolio optimization approaches primarily rely on stati…

View free PDFSource page
arxivcs.AIcs.CLcs.LGcs.MA2026-07-02

What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates

Arman Ghaffarizadeh, Danyal Mohaddes, Aliakbar Izadkhah, Shahriar Noroozizadeh

LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly to say. We study whether such social structure, without any explicit objective in the prompt, changes what an agent expresses pub…

View free PDFSource page