arxivcs.CLcs.AI2026-07-04
Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
Diego Cerda-Mardini, Sarath Chandar, Sreenath Madathil
LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language. For this role to be viable, models must produce identical verbal descriptions for identical i…