CORTEXA
← Browse

Andrew Mack

2 papers indexed

arxivcs.LGcs.AI2026-07-22

Scaling Interpretable Transformers with Parity Bottleneck Layers

Andrew Mack, Kraig Yuheng Tou, Mark Henry, Zhengxun Wu, Lauren Greenspan

Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are designed to recover such features post-hoc, but training models that are interpretable by construction…

View free PDFSource page