ChipsThe story, in brief

NVIDIA Introduces a 4-Bit Pretraining Methodology Using NVFP4, Validated on a 12B Hybrid Mamba-Transformer at 10T Token Horizon

10 trillion tokens trained in 4-bit. NVIDIA just proved you don't need full precision to scale.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

NVIDIA's NVFP4 microscaling format enables efficient large-scale pretraining without accuracy loss, directly reducing compute costs and power consumption for foundation model builders—a critical efficiency breakthrough in an era of GPU scarcity and rising training capex.

The key facts

5 to know
  1. NVFP4 microscaling format combines selective BF16 layers, 16×16 Random Hadamard Transforms, 2D weight scaling, and stochastic rounding

  2. 12B hybrid Mamba-Transformer trained on 10 trillion tokens—longest publicly documented 4-bit pretraining run

  3. Downstream accuracy: 62.58% (4-bit NVFP4) vs 62.62% (FP8 baseline) on MMLU-Pro, negligible divergence

  4. Methodology validated at production scale, not toy models

  5. Direct implication: lower GPU memory footprint, reduced training time, lower power draw per token

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: NVIDIA introduces a 4-bit pretraining methodology built around the NVFP4 microscaling format — combining selective BF16 layers, 16×16 Random Hadamard Transforms on Wgrad inputs, 2D weight scaling, and stochastic rounding on gradients — validated on a 12B hybrid Mamba-Transformer trained on 10…
Read original report
Back to today's editionMore chips news

Keep reading

Related stories

More from Chips