Preprint says confidence-guided decoding leads masked diffusion models into a shortcut that magnifies arithmetic errors
Masked diffusion models that order generation by confidence ignore long-range dependencies; the authors' own tests say training objectives amplify the error — worth tracking.
A preprint names the "confidence shortcut": masked diffusion models that order generation by confidence neglect long-range dependencies.
Masked diffusion models (MDMs) generate text by progressively unmasking tokens in any order, and could in principle reveal intermediate steps along logical dependencies. Authors Dueun Kim and Albert No report that standard decoding simply prioritizes high-confidence tokens, so in multi-digit addition the models predict higher-order digits without tracking carry chains.
The authors' own controlled pretraining shows confidence-guided ordering often selects suboptimal sequences, and confidence-aligned training objectives can raise addition error rates by an order of magnitude. The experimental code is open source; the findings come from the authors' own setup and await independent reproduction.
Sources:https://arxiv.org/abs/2605.29123https://github.com/jinha2536/mdm-arithmetic