Thursday, September 25, 2025
Top story
The Agent RaceOpenAI Blog
Measuring the performance of our models on real-world tasks
OpenAI's new GDPval evaluation framework measures model performance on economically valuable real-world tasks, not synthetic benchmarks. This shifts how the industry assesses AI capability beyond academic metrics—directly impacting how enterprises evaluate AI readiness for deployment.
New evaluation framework: GDPval
The briefs
OpenAI is expanding ChatGPT's enterprise surface area beyond chat—moving into team workflows, tool integrations, and compliance. This is the productization of AI as infrastructure, not just a consumer feature.
ChatGPT business plans now support shared projects