arxivcs.LGstat.ML2026-07-01
Prototype Language Models
Dan Ley, Giang Nguyen, Himabindu Lakkaraju, Julius Adebayo
Knowing which training examples drive outputs is fundamental to auditing, correcting, and understanding language models, yet for modern LLMs this remains expensive, approximate, and largely post-hoc. Standard language models generate tokens through a dense network pathway, causin…