FrontierThe story, in brief

Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput

2.03x throughput. NVIDIA's new Nemotron compress 75B model just changed the inference efficiency playbook.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

NVIDIA ships a production-ready compressed MoE variant that delivers massive throughput gains on B200/H100 hardware. For infrastructure teams, this is a direct cost-per-inference win that could reshape deployment economics.

The key facts

6 to know
  1. Nemotron-Labs-3-Puzzle-75B-A9B: compressed from 120.7B total / 12.8B active to 75.3B / 9.3B parameters

  2. 2.03x throughput improvement over Nemotron-3-Super on 8xB200 node

  3. 100 tokens/sec per user on single 8xB200 node

  4. 8x concurrency increase on single H100 (1 → 8 requests at 1M-token context)

  5. Compression method: iterative Puzzle alternating hardware-aware structural compression with knowledge distillation

  6. Published: July 9, 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.3B / 9.3B. On a single…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier