FrontierThe story, in brief

Google’s Gemma 4 open AI models use “speculative decoding” to get up to 3x faster - Ars Technica

3x faster. That's what Google's Gemma 4 just achieved—without cutting corners on quality.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Google's multi-token prediction drafters represent a meaningful inference optimization breakthrough for open models, directly impacting deployment economics and competitive positioning against closed-model inference speeds.

The key facts

6 to know
  1. Gemma 4 inference speed improvement: up to 3x faster

  2. Technology: speculative decoding / multi-token prediction (MTP) drafters

  3. Key claim: no quality loss with speed gains

  4. Model category: open-source, edge-optimized

  5. Use case emphasis: mobile and edge deployment

  6. Published: May 6, 2026

Go to the source

Reuters Technologynews.google.com

Publisher excerpt: Google’s Gemma 4 open AI models use “speculative decoding” to get up to 3x faster Ars Technica Accelerating Gemma 4: faster inference with multi-token prediction drafters blog.google Google AI Releases Multi-Token Prediction (MTP) Drafters for Gemma 4: Delivering Up to 3x Faster Inference Without…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier