CORTEXA
← Browse

Stefan Heimersheim

3 papers indexed

arxivcs.LG2026-07-06

Compressed Computation under $L^4$ Loss is likely Computation in Superposition

Francisco Ferreira da Silva, Stefan Heimersheim

Neural networks are thought to represent concepts as directions in their activation space, and superposition lets them encode more concepts than they have dimensions. It is natural to ask whether they can also compute more functions than they have neurons, i.e., perform computati…

View free PDFSource page
arxivcs.LG2026-07-01

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

Andrzej Szablewski, Gabriel Konar-Steenberg, Raffaello Fornasiere, Nikita Menon, Stefan Heimersheim

Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques. Current MOs are typically constructed via post-hoc supervised fine-tuning (SFT) on behavioural tran…

View free PDFSource page