arxivcs.LGcs.AIcs.CLstat.ML2026-06-26
What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs
Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath
Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs. Yet it remains unclear whether these explanations are sufficient, i.e., if they contain…