WorkJune 23, 2026via Apple Machine Learning
Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
Why it matters
Apple's research reveals that LLM evaluation panels suffer from correlated errors, meaning industry-standard multi-model judging produces far less independent information than assumed. This has immediate implications for how companies benchmark AI systems and make model selection decisions.
Key signals
- Panel of 9 frontier LLMs from 7 model families tested
- 9 judges effectively provide only ~2 independent votes of information
- ~75% of panel's nominal independence is lost to correlation
- Tested on three natural language inference datasets with 100 human annotations per item
- Published by Apple Machine Learning Research
- Challenges reliability of current LLM-as-a-judge evaluation methodology
- Only ~2 independent votes of informational value across 9 judges
- ~75% of panel's nominal independence is redundant
- Testing conducted on 3 NLI datasets with 100 human annotations per item
The hook
Nine judges, two votes. Apple researchers expose a critical flaw in how AI systems evaluate each other.
LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quantify how far their reliability falls short of the independent-voting ideal. Testing a panel of 9 frontier LLMs from 7 model families on three natural language inference datasets (each with 100 human annotations per item), we find that the 9 judges effectively provide only about 2 independent votes’ worth of information. Roughly three-quarters of the panel’s nominal independence…