rating-the-pitch-not-the-product-user-evaluations-of-llms-reflect-expectations-more-than-performance-03a6a43c·1 events·first seen Aliases: Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
A controlled experiment with 162 participants tested six LLMs across three tasks, manipulating whether users were told the model was better or worse than it actually was. Post-interaction ratings were predicted by whether the model met expectations (β≈0.47–0.50) and user confidence, not by task performance (β≈-0.01–0.11, n.s.). Actual output quality depended only on true model capability, not framing. The authors conclude that user-elicited evaluations—including preference data driving public leaderboards—measure expectation management at least as much as model quality.