arxivcs.CLcs.LG2026-06-26
Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution
Kevin Der, Harish Kamath, Ben Thompson
Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models. However, standard SAE architectures operate on individual token activations, meaning that the number of active features scales linearly with context length, and studying…