ChipsThe story, in brief

GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

Meta doubled training efficiency on its ads foundation model to 20-25% MFU while scaling 4x. Here's how they're optimizing the compute buildout.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Meta's GEM training post reveals concrete optimization patterns for LLM-scale workloads on thousands of GPUs—a data point in how frontier labs are pushing utilization economics as models grow. Practitioners scaling foundation models internally will extract technical details; enthusiasts tracking the compute buildout get evidence that efficiency gains are still available at scale.

The key facts

7 to know
  1. GEM (Generative Ads Recommendation Model) trains at LLM scale on several thousand latest-generation GPUs

  2. Doubled end-to-end training efficiency to 20-25% Model FLOPs Utilization (MFU)

  3. Scaled training FLOPs 4x while improving efficiency

  4. Foundation model powers ads recommendations across Instagram and Facebook

  5. Published by Meta Engineering on LLM-scale training optimization

  6. GEM trains at LLM scale on thousands of latest-generation GPUs

  7. Published August 2026 by Meta Engineering

Go to the source

Meta Engineeringengineering.fb.com

Publisher excerpt: Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end (E2E) training…
Read original report
Back to today's editionMore chips news

Keep reading

Related stories

More from Chips