arxivcs.LGcs.AIcs.CL2026-07-03
Individual Parameters in Weight-Sparse Transformers Appear Interpretable
Arnau Marin-Llobet, Stefan Heimersheim
A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does. Dominant circuit-finding approaches focus on a specific behavior and reverse-engineer the role of components on the associated sub-distribution. Howev…