NVIDIA Introduces a 4-Bit Pretraining Methodology Using NVFP4, Validated on a 12B Hybrid Mamba-Transformer at 10T Token Horizon
10 trillion tokens trained in 4-bit. NVIDIA just proved you don't need full precision to scale.

Why it matters
NVIDIA's NVFP4 microscaling format enables efficient large-scale pretraining without accuracy loss, directly reducing compute costs and power consumption for foundation model builders—a critical efficiency breakthrough in an era of GPU scarcity and rising training capex.
The key facts
5 to knowNVFP4 microscaling format combines selective BF16 layers, 16×16 Random Hadamard Transforms, 2D weight scaling, and stochastic rounding
12B hybrid Mamba-Transformer trained on 10 trillion tokens—longest publicly documented 4-bit pretraining run
Downstream accuracy: 62.58% (4-bit NVFP4) vs 62.62% (FP8 baseline) on MMLU-Pro, negligible divergence
Methodology validated at production scale, not toy models
Direct implication: lower GPU memory footprint, reduced training time, lower power draw per token
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: NVIDIA introduces a 4-bit pretraining methodology built around the NVFP4 microscaling format — combining selective BF16 layers, 16×16 Random Hadamard Transforms on Wgrad inputs, 2D weight scaling, and stochastic rounding on gradients — validated on a 12B hybrid Mamba-Transformer trained on 10…