Limits of Confidence in Diffusion
Apple researchers prove discrete diffusion has a fundamental flaw: per-position distributions can't match real token dependencies.

Why it matters
Apple's machine-learning team publishes research showing discrete diffusion models (used in image, audio, and text generation) have an inherent mathematical limitation: they cannot faithfully sample from distributions where tokens have dependencies. This affects remasking and uniform-state samplers, constraining what these models can learn and generate.
The key facts
10 to knowDiscrete diffusion writes multiple token positions per step from per-position distributions
A sampling step matches training distribution only when positions are conditionally independent given fixed tokens
No product of per-position distributions can match dependent groups
Applies to domains with inherent token dependencies: pixels, phonemes, words
Published by Apple Machine Learning Research, October 2026
Discrete diffusion (remasking, uniform-state samplers) writes multiple token positions per step
Step can only match training distribution when positions are conditionally independent given fixed tokens
No product of per-position distributions can match a dependent token group
Applies to domains with inherent dependencies: pixels, phonemes, words
Published by Apple Machine Learning Research, Oct 2, 2026
Go to the source
Apple Machine Learningmachinelearning.apple.com
Publisher excerpt: Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions. For domains of general interest (pixels, phonemes,…