arxivcs.CLcs.LG2026-07-01
A Mechanistic View of Authority Hierarchy in LLM Sycophancy
Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen
Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answers based on source credibility rather than evidence. We mechanistically investigate this phenomenon…