GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
Meta doubled training efficiency on its ads foundation model to 20-25% MFU while scaling 4x. Here's how they're optimizing the compute buildout.

Why it matters
Meta's GEM training post reveals concrete optimization patterns for LLM-scale workloads on thousands of GPUs—a data point in how frontier labs are pushing utilization economics as models grow. Practitioners scaling foundation models internally will extract technical details; enthusiasts tracking the compute buildout get evidence that efficiency gains are still available at scale.
The key facts
7 to knowGEM (Generative Ads Recommendation Model) trains at LLM scale on several thousand latest-generation GPUs
Doubled end-to-end training efficiency to 20-25% Model FLOPs Utilization (MFU)
Scaled training FLOPs 4x while improving efficiency
Foundation model powers ads recommendations across Instagram and Facebook
Published by Meta Engineering on LLM-scale training optimization
GEM trains at LLM scale on thousands of latest-generation GPUs
Published August 2026 by Meta Engineering
Go to the source
Meta Engineeringengineering.fb.com
Publisher excerpt: Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end (E2E) training…