Parallel decoding in diffusion models deviates 29× the noise floor
Confidence-ranked parallel token writing provably deviates from the training distribution; a synthetic task measured 29× the noise floor, real-world impact unquantified.
重要度局所的証拠E2 未複製執筆簡易
When discrete diffusion models write tokens in parallel by confidence ranking, any step writing two or more undetermined positions necessarily involves dependencies and cannot match the training distribution — Apple's Machine Learning Research group proved this in October, measuring on the synthetic task ScanAndAdd a total-variation distance of 29× the sampling-noise floor.
Before this, standard per-sample metrics read 1.0 and missed the deviation entirely, so samplers relying on parallel decoding had been drifting from the training distribution unnoticed.
The measurement was done by Apple's researchers: on ScanAndAdd the total-variation distance is 29× the noise floor while per-sample metrics read 1.0, meaning standard metrics cannot detect the deviation. The result applies directly to remasking and uniform-state samplers. The measurement so far covers only the authors' own synthetic task; the magnitude of the effect on real text and images remains unquantified, and no one has yet reproduced it.