arxivcs.LGcs.CL2026-07-01
Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance
Sinie van der Ben, Neele Roch, Anna Hedström, Mennatallah El-Assady
Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For these comparisons to be meaningful, scores must re…