FrontierAugust 11, 2026via The Decoder
Nvidia's open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence
Why it matters
Nvidia's open-weight release challenges the scaling doctrine by proving a 3.6B model can match much larger competitors on capability benchmarks while delivering industry-leading throughput. This signals a strategic pivot toward efficient inference and deployment at the edge — a capability bar shift practitioners building cost-sensitive AI systems will act on immediately.
Key signals
- Nemotron 3.5 Lightning: 3.6B active parameters
- Matches OpenAI's gpt-oss-120b on Intelligence Index despite being 4x smaller
- 670 tokens per second throughput (fastest in comparison)
- Open-weights release
- Efficiency-over-scale design philosophy
The hook
3.6B parameters. 670 tokens/sec. Nemotron 3.5 Lightning matches GPT-4 on reasoning — Nvidia's bet on efficiency over scale is reshaping the frontier.
Nvidia's Nemotron 3.5 Lightning is an open-weights model with just 3.6 billion active parameters that matches OpenAI's gpt-oss-120b on the Intelligence Index despite being four times smaller. At nearly 670 tokens per second, it's also the fastest model in the comparison, showing Nvidia is betting on…