Google’s Gemma 4 open AI models use “speculative decoding” to get up to 3x faster - Ars Technica
3x faster. That's what Google's Gemma 4 just achieved—without cutting corners on quality.

Why it matters
Google's multi-token prediction drafters represent a meaningful inference optimization breakthrough for open models, directly impacting deployment economics and competitive positioning against closed-model inference speeds.
The key facts
6 to knowGemma 4 inference speed improvement: up to 3x faster
Technology: speculative decoding / multi-token prediction (MTP) drafters
Key claim: no quality loss with speed gains
Model category: open-source, edge-optimized
Use case emphasis: mobile and edge deployment
Published: May 6, 2026
Go to the source
Reuters Technologynews.google.com
Publisher excerpt: Google’s Gemma 4 open AI models use “speculative decoding” to get up to 3x faster Ars Technica Accelerating Gemma 4: faster inference with multi-token prediction drafters blog.google Google AI Releases Multi-Token Prediction (MTP) Drafters for Gemma 4: Delivering Up to 3x Faster Inference Without…