FrontierThe story, in brief

NVIDIA Releases Nemotron-Labs-TwoTower: an Open-Weight Diffusion Language Model Built on a Frozen Autoregressive Nemotron-3-Nano-30B-A3B Backbone

NVIDIA just open-sourced a diffusion language model that sidesteps the autoregressive bottleneck. Here's why throughput matters more than you think.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

NVIDIA's Nemotron-Labs-TwoTower challenges the autoregressive paradigm by using discrete diffusion to parallelize token generation, directly addressing a fundamental inference constraint that impacts production deployment economics.

The key facts

7 to know
  1. Model: Nemotron-Labs-TwoTower (diffusion-based language model)

  2. Built on frozen Nemotron-3-Nano-30B-A3B autoregressive backbone

  3. Released as open-weight under NVIDIA Nemotron Open Model License

  4. Target: throughput bottleneck in text generation

  5. Architecture: discrete diffusion vs. serial autoregressive decoding

  6. Addresses serial token-by-token generation constraint

  7. Published Jul 01 2026 on MarkTechPost

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: NVIDIA has released Nemotron-Labs-TwoTower, a diffusion language model built on a pretrained autoregressive backbone. It ships as open weights under the NVIDIA Nemotron Open Model License. The release targets a throughput bottleneck in text generation. Autoregressive (AR) models decode one token at…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier