Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
Nine judges, two votes. Apple researchers expose a critical flaw in how AI systems evaluate each other.

Why it matters
Apple's research reveals that LLM evaluation panels suffer from correlated errors, meaning industry-standard multi-model judging produces far less independent information than assumed. This has immediate implications for how companies benchmark AI systems and make model selection decisions.
The key facts
9 to knowPanel of 9 frontier LLMs from 7 model families tested
9 judges effectively provide only ~2 independent votes of information
~75% of panel's nominal independence is lost to correlation
Tested on three natural language inference datasets with 100 human annotations per item
Published by Apple Machine Learning Research
Challenges reliability of current LLM-as-a-judge evaluation methodology
Only ~2 independent votes of informational value across 9 judges
~75% of panel's nominal independence is redundant
Testing conducted on 3 NLI datasets with 100 human annotations per item
Go to the source
Apple Machine Learningmachinelearning.apple.com
Publisher excerpt: LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quantify how far their reliability falls short of the independent-voting ideal.…