arxivcs.AIcs.CR2026-07-11
When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control
We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior. This question is difficult because apparent success can arise from weak interventions, mismatched baselines, model robustness, or degenerate outputs that automated sa…