DiffusionGemma: 4x faster text generation
4x faster. Google DeepMind just rewrote how text generation works.

Why it matters
DiffusionGemma represents a fundamental shift in inference speed through diffusion-based decoding, directly challenging the speed-vs-quality tradeoff that has defined LLM deployment economics. This could reshape inference cost calculations across the industry.
The key facts
5 to knowDiffusionGemma achieves 4x speedup in text generation
Uses diffusion-based approach (non-autoregressive or hybrid decoding paradigm)
Published by Google DeepMind
Published June 10, 2026
Addresses inference latency — a critical bottleneck in LLM production
Go to the source
Google DeepMind Blogdeepmind.google