arxivcs.CLcs.CYcs.HC2026-07-06
Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
Robert Morabito, Tyler McDonald, Charitra Viswanath, Angel Hsing-Chi Hwang, Susanne Gaube, Jad Kabbara, et al.
Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model; the other, an older, weaker model. They walk away with markedly different ratings of its usefulness and intelligence, yet they used the same model. In a controlled study, 162 pa…