Accounting for cognitive bias in human evaluation of large language models
Everyone is focused on model benchmarks. Nobody is talking about the cognitive bias poisoning human evaluations of LLMs.

Why it matters
Amazon Science's framework challenges how we measure LLM quality. If human evaluation is systematically biased, every benchmark we trust might be wrong—forcing AI teams to rethink evaluation methodology.
The key facts
8 to knowPosition paper presented at ACL (Association for Computational Linguistics)
Focus on cognitive bias in human evaluation of LLMs
Proposes framework for more accurate evaluation methodology
Amazon Science research
Published September 2024
Position paper presented at ACL (major NLP conference)
Amazon Science authorship suggests internal deployment relevance
Addresses gap between evaluation metrics and real-world performance
Go to the source
Amazon Scienceamazon.science
Publisher excerpt: A position paper presented at ACL proposes a framework for more-accurate human evaluation of LLMs.

