ToolsThe story, in brief

OpenAI Releases GPT-Realtime-2.1 and GPT-Realtime-2.1-mini for Low-Latency Voice Agents in the API

25% latency cut. OpenAI's new Realtime-2.1 models just made voice agents production-ready for enterprises.

Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI agents and the coordination of work.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenAI is shipping lower-latency voice models with improved reasoning capabilities, directly enabling real-time AI agent deployment at scale. This moves voice from demo to deployment—critical for enterprises building conversational workflows.

The key facts

6 to know
  1. Two new Realtime models: GPT-Realtime-2.1 and GPT-Realtime-2.1-mini

  2. P95 latency reduced by at least 25% through improved caching

  3. GPT-Realtime-2.1-mini includes reasoning capability

  4. Pricing unchanged for mini variant

  5. WebRTC connection support for API integration

  6. Published July 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: OpenAI added two Realtime models to its API. GPT-Realtime-2.1-mini is a mini reasoning model for voice, priced like the earlier gpt-realtime-mini. OpenAI also cut p95 latency by at least 25% through improved caching. Here is what changed, how pricing compares, and how to connect over WebRTC.
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools