arxivcs.LGcs.AI2026-07-17
Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent
Sriram Balasubramanian, Soheil Feizi
Interpretability methods for neural network activations span a wide cost spectrum, from cheap, training-free techniques (such as linear probes, PCA, SVD) to more expensive training-based ones (such as SAEs and activation oracles). Training-based methods are typically more powerfu…