WorkThe story, in brief

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

Nine judges, two votes. Apple researchers expose a critical flaw in how AI systems evaluate each other.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Apple's research reveals that LLM evaluation panels suffer from correlated errors, meaning industry-standard multi-model judging produces far less independent information than assumed. This has immediate implications for how companies benchmark AI systems and make model selection decisions.

The key facts

9 to know
  1. Panel of 9 frontier LLMs from 7 model families tested

  2. 9 judges effectively provide only ~2 independent votes of information

  3. ~75% of panel's nominal independence is lost to correlation

  4. Tested on three natural language inference datasets with 100 human annotations per item

  5. Published by Apple Machine Learning Research

  6. Challenges reliability of current LLM-as-a-judge evaluation methodology

  7. Only ~2 independent votes of informational value across 9 judges

  8. ~75% of panel's nominal independence is redundant

  9. Testing conducted on 3 NLI datasets with 100 human annotations per item

Go to the source

Apple Machine Learningmachinelearning.apple.com

Publisher excerpt: LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quantify how far their reliability falls short of the independent-voting ideal.…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work