arxivcs.CLcs.AIcs.LG2026-06-26
VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring
Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah, Martha Lewis
Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post hoc rather than directly connected to the Transformer's token vocabulary. We introduce Vocabulary-Aligned Sparse Autoencoder (VASAE), a meth…