FrontierAugust 9, 2026via The Decoder

Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model

Why it matters

DiffusionGemma demonstrates a new path to faster inference by adapting existing models rather than training from scratch, with material speed gains but quality tradeoffs that matter for reasoning workloads.

Key signals

  • DiffusionGemma adapted from Gemma 4 using <10% of original training budget
  • Generates 256 tokens in parallel (vs. one at a time in autoregressive)
  • Throughput: ~1,500 tokens per second
  • Quality trails original autoregressive model, especially on reasoning tasks
  • Model adaptation/retrofitting approach as alternative to from-scratch training

The hook

Google retrofitted Gemma into a diffusion model using <10% of training budget. Parallel generation hits 1.5K tokens/sec — but reasoning quality drops.

Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second. Quality still trails t

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.