Fit More and Train Faster With ZeRO via DeepSpeed and FairScale
Not a pilot. Hugging Face just made large-scale AI model training 3x more efficient.

Why it matters
ZeRO optimization via DeepSpeed and FairScale addresses a critical enterprise bottleneck: the cost and computational overhead of training large language models. This enables smaller teams and organizations to train models previously requiring massive infrastructure budgets.
The key facts
10 to knowZeRO optimization technology released
DeepSpeed and FairScale integration
Enables training of larger models with same hardware resources
Published January 19, 2021 (Hugging Face blog)
Direct impact on model training efficiency and accessibility
ZeRO optimization reduces training memory footprint by up to 10x
Enables training of larger models on fewer GPUs
Collaboration between Microsoft (DeepSpeed) and Meta/Facebook (FairScale)
Published January 19, 2021 - technical infrastructure release
Addresses core constraint: memory efficiency in distributed training
Go to the source
Hugging Face Bloghuggingface.co