FrontierThe story, in brief

Thinking Machines Lab ships its first model and argues interactivity is what OpenAI gets wrong about voice

Mira Murati's new lab just shipped a multimodal model that processes audio, video, and text in 200ms chunks—and it's directly challenging OpenAI's voice AI approach.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Thinking Machines Lab is entering the competitive multimodal voice AI space with a technical differentiation (parallel chunk processing) that positions it against OpenAI's GPT Realtime and Google's Gemini Live. This signals a new entrant attacking a core product category where incumbents are considered weak on interactivity.

The key facts

6 to know
  1. Mira Murati founded Thinking Machines Lab

  2. First model processes audio, video, and text in parallel

  3. 200-millisecond processing latency

  4. Direct competitive positioning against OpenAI GPT Realtime 2 and Google Gemini Live

  5. Focus on interactivity as differentiator vs. question-and-answer model limitations

  6. Multimodal capability (audio, video, text)

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Mira Murati's start-up presents its first AI model and aims to free voice AI from the question-and-answer model. The model processes audio, video and text in 200-millisecond chunks in parallel and aims to beat OpenAI's GPT Realtime 2 and Google's Gemini Live in terms of interaction quality.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier