FrontierThe story, in brief

Nvidia's open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence

3.6B parameters. 670 tokens/sec. Nemotron 3.5 Lightning matches GPT-4 on reasoning — Nvidia's bet on efficiency over scale is reshaping the frontier.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Nvidia's open-weight release challenges the scaling doctrine by proving a 3.6B model can match much larger competitors on capability benchmarks while delivering industry-leading throughput. This signals a strategic pivot toward efficient inference and deployment at the edge — a capability bar shift practitioners building cost-sensitive AI systems will act on immediately.

The key facts

5 to know
  1. Nemotron 3.5 Lightning: 3.6B active parameters

  2. Matches OpenAI's gpt-oss-120b on Intelligence Index despite being 4x smaller

  3. 670 tokens per second throughput (fastest in comparison)

  4. Open-weights release

  5. Efficiency-over-scale design philosophy

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Nvidia's Nemotron 3.5 Lightning is an open-weights model with just 3.6 billion active parameters that matches OpenAI's gpt-oss-120b on the Intelligence Index despite being four times smaller. At nearly 670 tokens per second, it's also the fastest model in the comparison, showing Nvidia is betting…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier