arxivcs.LG2026-07-08
Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning
This article offers a comprehensive overview of mechanistic interpretability, an emerging field that seeks to reverse-engineer the internal algorithms of modern neural networks. While traditional explainable AI methods often stop at surface-level input-output correlations, this a…