FrontierThe story, in brief

OpenAI Releases Three Realtime Audio Models: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in the Realtime API

Three new audio models. OpenAI just made voice agents and real-time translation table stakes for every developer.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenAI expands its model portfolio with three specialized audio capabilities (reasoning, translation, transcription), raising the bar for competitors and enabling a new class of voice-first applications that developers can build immediately.

The key facts

5 to know
  1. Three new models: GPT-Realtime-2, GPT-Realtime-Translate, GPT-Realtime-Whisper

  2. GPT-Realtime-Translate supports 70+ languages

  3. Models designed for live voice, reasoning agents, and streaming transcription

  4. Available via Realtime API

  5. Published May 8, 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Three purpose-built audio models expand what developers can build with live voice: reasoning agents, speech translation across 70+ languages, and streaming transcription. The post OpenAI Releases Three Realtime Audio Models: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in the…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier