FrontierThe story, in brief

AI benchmarks systematically ignore how humans disagree, Google study finds

Three to five human raters. That's what most AI benchmarks use - and Google's new study shows it's not nearly enough.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

This research exposes a fundamental flaw in how AI models are evaluated, potentially invalidating benchmark results that companies use to make critical AI deployment decisions.

The key facts

4 to know
  1. Standard benchmarks use only 3-5 human raters per test example

  2. Current rating methodology produces unreliable results

  3. Annotation budget allocation matters as much as total budget

  4. Human disagreement is systematically ignored in current benchmarks

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: A Google study finds that the standard three to five human raters per test example often aren't enough for reliable AI benchmarks, and that splitting your annotation budget the right way matters just as much as the budget itself.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier