FACTS Benchmark Suite: Systematically evaluating the factuality of large language models
Google DeepMind just released a new way to measure what LLMs actually get wrong. Here's why your model's accuracy claims might be overstated.

Why it matters
Google DeepMind's FACTS Benchmark Suite introduces systematic evaluation methodology for LLM factuality—a critical capability gap that impacts production deployment decisions and competitive model positioning.
The key facts
5 to knowGoogle DeepMind released FACTS Benchmark Suite
Focuses on systematic factuality evaluation of LLMs
Published December 9, 2025
Addresses gap in standardized factuality measurement across models
Relevant to model comparison and capability assessment
Go to the source
Google DeepMind Blogdeepmind.google
Publisher excerpt: Systematically evaluating the factuality of large language models with the FACTS Benchmark Suite.