CORTEXA
← Browse
arxivcs.LG2026-07-08

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning

Pranav Sawant, Jakub Krejčí

This article offers a comprehensive overview of mechanistic interpretability, an emerging field that seeks to reverse-engineer the internal algorithms of modern neural networks. While traditional explainable AI methods often stop at surface-level input-output correlations, this approach directly addresses the opaque "black box" nature of machine learning models, which is essential for ensuring safety and auditability in high-stakes deployments. The paper provides a detailed examination of Transformer circuit analysis, exploring how internal components like the residual stream, attention mechanisms, and induction heads drive complex tasks and in-context learning. It subsequently tackles the core challenge of superposition and polysemanticity, demonstrating how tools like Sparse Autoencoders (SAEs) and transcoders can decompose tangled network activations into distinct, human-interpretable features. Furthermore, the paper explores methods for actively controlling and modifying model behavior through steering vectors and causal interventions. Finally, it connects these mechanistic insights with neurosymbolic AI frameworks designed to translate neural representations into explicit, executable logical rules.

View free PDFSource page

Related papers

arxivquant-phcond-mat.dis-nncond-mat.str-elcs.AIcs.LG2026-07-01

Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders

Zihao Qi, Christopher Earls

Neural Quantum States (NQS) are a remarkably expressive class of variational ansätze for quantum many-body wavefunctions, yet little is understood about their internal mechanisms: trained on variational objectives alone, how do NQS accurately capture physical observables that the…

View free PDFSource page
arxivcs.LG2026-06-29

Improved Predictive Performance and Interpretability for Mesomorphic Neural Networks Using Local Fidelity Regularization

Hugo L. Hammer, Vajira Thambawita, Kristoffer Herland Hellton, Pål Halvorsen

Interpretable Mesomorphic Neural Networks (IMNs) offer a promising framework that combines the predictive power of deep neural networks with the interpretability of linear models. However, the original formulation lacks safeguards to ensure that the learned interpretations are in…

View free PDFSource page
arxivnucl-thcs.LG2026-06-26

Bridging Ab Initio Symmetries and Global Nuclear Masses with Interpretable Neural Networks

Phong Dang, Evander Espinoza, Xiaoliang Wan, Michela Negro, Jerry P. Draayer, Feng Pan, et al.

Ab initio modeling has established Wigner's SU(4) and Elliott's SU(3) as dominant symmetries of the nuclear force in light and intermediate-mass nuclei. We ask whether they also govern nuclear binding across the entire chart. Our aim is not high-precision prediction but physical…

View free PDFSource page
arxivcs.LGphysics.flu-dyn2026-07-13

A multi-scale feature enhanced graph neural network for fluid dynamics prediction in complex geometries

Li Xiao, Tianyu Li, Yiye Zou, Mingjie Zhang, Xiaogangd Deng

Industrial design in fields such as vehicle and aerospace engineering often relies on large-scale numerical simulations to evaluate fluid dynamics performance, which can incur substantial computational costs. Deep neural networks have shown promise in improving simulation efficie…

View free PDFSource page
arxivcs.LGcs.AIcs.IT2026-07-02

Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability

Rodrigo Mendoza-Smith

Sparse autoencoders (SAEs) decompose internal activations of neural networks into sparse linear combinations of learned features by fitting an overcomplete dictionary $\mathbf{W}\in\mathbb{R}^{m\times n}$ with $m<n$, and inferring a sparse code $\mathbf{x}\in\mathbb{R}^n$ from $\…

View free PDFSource page
arxivcs.LG2026-06-25

Explaining Temporal Graph Neural Networks via Feature-induced Information Flow

Ping Xiong, Thomas Schnake, Klaus-Robert Müller, Shinichi Nakajima

Event-based Temporal Graph Neural Networks (ETGNNs) have demonstrated strong performance across a wide range of applications, including social network analysis, epidemic tracing, recommender systems, and political event forecasting. However, their increasing complexity poses sign…

View free PDFSource page