FrontierAugust 11, 2026via The Decoder

Nvidia's open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence

Why it matters

Nvidia's open-weight release challenges the scaling doctrine by proving a 3.6B model can match much larger competitors on capability benchmarks while delivering industry-leading throughput. This signals a strategic pivot toward efficient inference and deployment at the edge — a capability bar shift practitioners building cost-sensitive AI systems will act on immediately.

Key signals

  • Nemotron 3.5 Lightning: 3.6B active parameters
  • Matches OpenAI's gpt-oss-120b on Intelligence Index despite being 4x smaller
  • 670 tokens per second throughput (fastest in comparison)
  • Open-weights release
  • Efficiency-over-scale design philosophy

The hook

3.6B parameters. 670 tokens/sec. Nemotron 3.5 Lightning matches GPT-4 on reasoning — Nvidia's bet on efficiency over scale is reshaping the frontier.

Nvidia's Nemotron 3.5 Lightning is an open-weights model with just 3.6 billion active parameters that matches OpenAI's gpt-oss-120b on the Intelligence Index despite being four times smaller. At nearly 670 tokens per second, it's also the fastest model in the comparison, showing Nvidia is betting on

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.