CORTEXA
← Browse

Sinie van der Ben

2 papers indexed

arxivcs.LGcs.CL2026-07-01

Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance

Sinie van der Ben, Neele Roch, Anna Hedström, Mennatallah El-Assady

Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For these comparisons to be meaningful, scores must re…

View free PDFSource page
arxivcs.CLcs.AI2026-06-25

Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

Sinie van der Ben, Raphaël Baur, Yannick Metz, Mennatallah El-Assady

Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirroring human psychological structure. We test the generality of these findings in two open-weight mod…

View free PDFSource page