arxivcs.AI2026-06-27
Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions
Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. In this work, we propose a mechanistic interpretability appro…