FrontierThe story, in brief

NVIDIA AI Releases Nemotron-Labs-Diffusion: A Tri-Mode Language Model with 6× Tokens Per Forward Over Qwen3-8B

6×. That's how many more tokens NVIDIA's new Nemotron processes per forward pass versus Qwen3-8B.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

NVIDIA released a tri-mode language model architecture that unifies autoregressive, diffusion-based parallel, and self-speculation decoding in a single model family. This directly addresses inference throughput bottlenecks—a critical competitive dimension as enterprises optimize cost-per-token and latency.

The key facts

7 to know
  1. Model family: Nemotron-Labs-Diffusion

  2. Parameter sizes: 3B, 8B, 14B

  3. Three decoding modes: autoregressive (AR), diffusion-based parallel, self-speculation

  4. Performance claim: 6× tokens per forward pass vs Qwen3-8B

  5. Variants: base, instruct, vision-language

  6. Released by: NVIDIA AI research

  7. Published: May 20, 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: NVIDIA researchers have released Nemotron-Labs-Diffusion, a language model family that unifies three decoding modes in one architecture. The model supports autoregressive (AR) decoding, diffusion-based parallel decoding, and self-speculation decoding. It is available in 3B, 8B, and 14B parameter…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier