Introducing SimpleQA
OpenAI just released a new factuality benchmark. Here's why it matters for your model evals.

Why it matters
OpenAI introduced SimpleQA, a factuality benchmark designed to measure language models' ability to answer short, fact-seeking questions accurately. This is a critical evaluation tool for assessing model reliability in production use cases where hallucination and factual errors carry real cost.
The key facts
5 to knowSimpleQA is a factuality benchmark by OpenAI
Measures language model ability to answer short, fact-seeking questions
Released October 30, 2024
Addresses key evaluation gap: factual accuracy in closed-domain QA
Relevant for production AI deployments requiring high factuality standards
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.