FrontierThe story, in brief

NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput

2.03x throughput. NVIDIA's new Nemotron-Labs compresses a 120B model to 75B without losing performance—and servers can now handle 8x more concurrent requests on H100s.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

NVIDIA is shipping production-grade model compression that solves the inference efficiency problem at scale. For operators, this means 8x better concurrency on existing hardware—a direct path to lower latency and higher utilization without new capex.

The key facts

7 to know
  1. Nemotron-Labs-3-Puzzle-75B-A9B: compressed from 120.7B → 75.3B total parameters

  2. Active parameters: 12.8B → 9.3B (25% reduction in active compute)

  3. 2.03x total throughput on 8xB200 node at matched user throughput (100 tok/s per user)

  4. H100 concurrency: 1 request → 8 requests at 1M-token concurrency

  5. Compression method: iterative Puzzle (hardware-aware structural compression + knowledge distillation recovery)

  6. Released by NVIDIA Labs

  7. Published: July 9, 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.3B / 9.3B. On a single…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier