Measuring the performance of our models on real-world tasks
OpenAI just benchmarked AI across 44 real jobs. Here's what it means for your workforce.

Why it matters
OpenAI's new GDPval evaluation framework measures model performance on economically valuable real-world tasks, not synthetic benchmarks. This shifts how the industry assesses AI capability beyond academic metrics—directly impacting how enterprises evaluate AI readiness for deployment.
The key facts
5 to knowNew evaluation framework: GDPval
Scope: 44 occupations
Focus: Real-world economically valuable tasks
Departure from synthetic benchmarks
Published: September 25, 2025
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: OpenAI introduces GDPval, a new evaluation that measures model performance on real-world economically valuable tasks across 44 occupations.