Speculative Decoding Slowed by Attacks
New research claims adversarial suffixes can lower draft acceptance rates, making speculative decoding slower and more expensive than autoregression.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Speculative decoding, a mainstream acceleration technique, faces a risk of being maliciously slowed down.
Researchers propose "Speculative Rejection Attacks" (SRAs), which append adversarial suffixes to user inputs to force greater disagreement between draft and target models. This reduces the number of accepted tokens per cycle, requiring more forward passes from the target model.
The paper notes that in some cases, these attacks make inference slower than standard autoregressive decoding, significantly increasing computational costs. While regularization can restore output quality, it sacrifices most of the attack's effectiveness.
This finding identifies the draft-target interaction as a realistic attack surface. The results are currently preprint-only and have not been independently reproduced in large-scale production environments.