Thinking Machines Lab ships its first model and argues interactivity is what OpenAI gets wrong about voice
Mira Murati's new lab just shipped a multimodal model that processes audio, video, and text in 200ms chunks—and it's directly challenging OpenAI's voice AI approach.

Why it matters
Thinking Machines Lab is entering the competitive multimodal voice AI space with a technical differentiation (parallel chunk processing) that positions it against OpenAI's GPT Realtime and Google's Gemini Live. This signals a new entrant attacking a core product category where incumbents are considered weak on interactivity.
The key facts
6 to knowMira Murati founded Thinking Machines Lab
First model processes audio, video, and text in parallel
200-millisecond processing latency
Direct competitive positioning against OpenAI GPT Realtime 2 and Google Gemini Live
Focus on interactivity as differentiator vs. question-and-answer model limitations
Multimodal capability (audio, video, text)
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Mira Murati's start-up presents its first AI model and aims to free voice AI from the question-and-answer model. The model processes audio, video and text in 200-millisecond chunks in parallel and aims to beat OpenAI's GPT Realtime 2 and Google's Gemini Live in terms of interaction quality.