FrontierThe story, in brief

Taming Outlier Tokens in Diffusion Transformers

Apple researchers identify a hidden inefficiency in Diffusion Transformers that could reshape how image-generation models are built.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Apple's research surfaces a technical problem (outlier tokens) in modern DiT architectures used for image generation. Understanding and fixing this could improve efficiency and quality — actionable for practitioners building or optimizing generative models.

The key facts

11 to know
  1. Outlier tokens identified in both encoder and denoiser of RAE-DiT pipelines

  2. Phenomenon appears especially in intermediate layers of DiTs

  3. Prior ViT work on outliers extended to generative model context for first time

  4. Published by Apple ML research team

  5. Focus on Representation Autoencoder (RAE) DiT architecture

  6. Study addresses underexplored role of outliers in generative models

  7. Outlier tokens appear in both encoder and denoiser of RAE-DiT pipelines

  8. Phenomenon especially pronounced in intermediate layers of DiTs

  9. Pretrained ViT encoders produce outlier representations that carry limited local information

  10. Research from Apple Machine Learning (published August 2026)

  11. Addresses inefficiency in modern image generation architectures

Go to the source

Apple Machine Learningmachinelearning.apple.com

Publisher excerpt: We study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention while carrying limited local information, but their role in generative models…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier