Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning
Speech-native reasoning just jumped 50 points. Kyutai's Voice of Reason solves spoken math end-to-end—no transcription, no text LLM—with RL.

Why it matters
A new capability frontier: models that reason directly in speech without transcription bottlenecks. This changes how we think about multimodal reasoning and what's possible with open-weight releases at scale.
The key facts
6 to knowKyutai released Voice of Reason: open-weight speech-to-speech models on Hugging Face
Built on GLM-4-Voice-9B base
Spoken GSM8K accuracy: 27.3% → 77.1% via supervised fine-tuning + reinforcement learning
No transcription step; no text LLM in the pipeline
Both checkpoints run on single H100
Demonstrates speech-native reasoning as a capability, not speech-as-input-to-text-LLM
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Kyutai has released Voice of Reason, 2 open-weight speech-to-speech models built on GLM-4-Voice-9B. Supervised fine-tuning and reinforcement learning lift spoken GSM8K accuracy from 27.3% to 77.1%. There is no transcription step and no text LLM in the loop. Both checkpoints are on Hugging Face and…