OpenAI Releases GPT-Realtime-2.1 and GPT-Realtime-2.1-mini for Low-Latency Voice Agents in the API
25% latency cut. OpenAI's new Realtime-2.1 models just made voice agents production-ready for enterprises.

Why it matters
OpenAI is shipping lower-latency voice models with improved reasoning capabilities, directly enabling real-time AI agent deployment at scale. This moves voice from demo to deployment—critical for enterprises building conversational workflows.
The key facts
6 to knowTwo new Realtime models: GPT-Realtime-2.1 and GPT-Realtime-2.1-mini
P95 latency reduced by at least 25% through improved caching
GPT-Realtime-2.1-mini includes reasoning capability
Pricing unchanged for mini variant
WebRTC connection support for API integration
Published July 2026
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: OpenAI added two Realtime models to its API. GPT-Realtime-2.1-mini is a mini reasoning model for voice, priced like the earlier gpt-realtime-mini. OpenAI also cut p95 latency by at least 25% through improved caching. Here is what changed, how pricing compares, and how to connect over WebRTC.