arxivcs.CLcs.AIcs.HCcs.LG2026-07-10
A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss
Kwan Soo Shin, In Seok Kang, Yunkyung Min, Munho Lee
Whether a language model behaves as it claims is a judgement on which independent human raters cannot agree (Fleiss kappa = 0.074). We show that a small, purpose-built instrument does better. A linear read-out of the frozen representation of a from-scratch 146-million-parameter a…