Apple research shows parallel decoding in diffusion models mismatches the training distribution
Confidence-ranked parallel token writing provably deviates from the training distribution; a synthetic task measured 29× the noise floor, real-world impact unquantified.
Apple's Machine Learning Research group published a paper in October proving that when discrete diffusion models write tokens in parallel by confidence ranking, any step writing two or more undetermined positions necessarily involves dependencies and cannot match the training distribution.
On ScanAndAdd, a synthetic task with a closed-form joint distribution, the paper measures the generated distribution at 29× the sampling-noise floor in total variation, while per-sample metrics read 1.0 — meaning standard metrics miss the deviation entirely.
The result applies directly to remasking and uniform-state samplers. The measurement so far covers only the authors' own synthetic task; the magnitude of the effect on real text and images remains unquantified.
Sources:https://machinelearning.apple.com/research/limits-confidence-diffusionhttps://arxiv.org/abs/2609.20581