arxivcs.CLcs.AI2026-07-14
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
Winston Zeng, Ali Emami, Jinho D. Choi
What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers o…