WorkJune 23, 2026via Apple Machine Learning

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

Why it matters

Apple's research reveals that LLM evaluation panels suffer from correlated errors, meaning industry-standard multi-model judging produces far less independent information than assumed. This has immediate implications for how companies benchmark AI systems and make model selection decisions.

Key signals

  • Panel of 9 frontier LLMs from 7 model families tested
  • 9 judges effectively provide only ~2 independent votes of information
  • ~75% of panel's nominal independence is lost to correlation
  • Tested on three natural language inference datasets with 100 human annotations per item
  • Published by Apple Machine Learning Research
  • Challenges reliability of current LLM-as-a-judge evaluation methodology
  • Only ~2 independent votes of informational value across 9 judges
  • ~75% of panel's nominal independence is redundant
  • Testing conducted on 3 NLI datasets with 100 human annotations per item

The hook

Nine judges, two votes. Apple researchers expose a critical flaw in how AI systems evaluate each other.

LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quantify how far their reliability falls short of the independent-voting ideal. Testing a panel of 9 frontier LLMs from 7 model families on three natural language inference datasets (each with 100 human annotations per item), we find that the 9 judges effectively provide only about 2 independent votes’ worth of information. Roughly three-quarters of the panel’s nominal independence…

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels | KeyNews.AI