Thursday, September 25, 2025

Start of archive·May 16

Top story

The Agent RaceOpenAI Blog

Measuring the performance of our models on real-world tasks

OpenAI's new GDPval evaluation framework measures model performance on economically valuable real-world tasks, not synthetic benchmarks. This shifts how the industry assesses AI capability beyond academic metrics—directly impacting how enterprises evaluate AI readiness for deployment.

New evaluation framework: GDPval

The briefs

OpenAI is expanding ChatGPT's enterprise surface area beyond chat—moving into team workflows, tool integrations, and compliance. This is the productization of AI as infrastructure, not just a consumer feature.

ChatGPT business plans now support shared projects