Tuesday, December 9, 2025

Start of archive·May 16

Top story

The Agent RaceGoogle DeepMind Blog

FACTS Benchmark Suite: Systematically evaluating the factuality of large language models

Google DeepMind's FACTS Benchmark Suite introduces systematic evaluation methodology for LLM factuality—a critical capability gap that impacts production deployment decisions and competitive model positioning.

Google DeepMind released FACTS Benchmark Suite

The briefs

OpenAI's co-founding of an industry standards body (Agentic AI Foundation under Linux Foundation) signals a strategic pivot toward interoperability and governance in agentic systems. This is a policy/standards play that shapes how the industry will build, deploy, and regulate autonomous agents—critical context for leaders deciding between proprietary vs. open-stack agent architectures.

OpenAI co-founds Agentic AI Foundation under Linux Foundation

Scout24 is deploying GPT-5 at the product layer to fundamentally change how users discover properties—moving from keyword search to guided, conversational discovery. This is a live example of how enterprise teams are building agent-like workflows into consumer-facing applications.

Scout24 deployed GPT-5 powered conversational assistant

OpenAI is moving beyond models into workforce development and credentialing, creating a new revenue stream while positioning itself as the de facto standard for AI skills training in enterprise.

OpenAI launches first official certification courses