The Agent RaceJuly 17, 2026via MarkTechPost

NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEB

Why it matters

NVIDIA is competing directly in the embedding model space with a top-ranked open release, signaling a strategic push beyond inference hardware into the foundation model stack. The 1B variant via distillation shows NVIDIA is optimizing for efficiency and edge deployment.

Key signals

  • Nemotron-3-Embed-8B ranks #1 on RTEB at 78.46 average NDCG@10
  • Three open checkpoints released: 8B, 1B (BF16), and 1B (NVFP4)
  • 1B model derived from ModelOpt NAS pruning + COS+MSE distillation
  • NVFP4 format retains 99%+ BF16 accuracy at up to 2x Blackwell throughput
  • All models support 32,768-token context windows under OpenMDW-1.1
  • Release date: July 15-16, 2026

The hook

#1 on RTEB. NVIDIA's Nemotron-3-Embed-8B just dethroned the field—and it's open-source.

NVIDIA released Nemotron 3 Embed on July 15 and 16, 2026. The collection has three open checkpoints: Nemotron-3-Embed-8B-BF16, Nemotron-3-Embed-1B-BF16, and Nemotron-3-Embed-1B-NVFP4. The 8B ranks #1 on RTEB at 78.46 average NDCG@10. The 1B came from ModelOpt NAS pruning plus COS+MSE distillation from the 8B teacher. NVFP4 retains 99%+ of BF16 retrieval accuracy at up to 2x Blackwell throughput. All three run 32,768-token inputs under OpenMDW-1.1.

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.