Layered Rounding Cuts Low-Precision Error
A study shows that using stochastic rounding for MLP layers and round-to-nearest for the output head reduces DistilGPT-2 perplexity degradation to 1.10x at 6-bit precision, compared to 2.21x for uniform round-to-nearest.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
At 6-bit virtual precision, a mixed rounding strategy keeps DistilGPT-2 perplexity degradation within 1.10x of the full-precision reference.
Previous approaches often applied either stochastic rounding (SR) or round-to-nearest (RN) uniformly across the network. This research reveals that different layers respond oppositely to rounding noise: MLP layer errors cause uniform logit shifts, which Softmax is invariant to, making SR's variance penalty negligible there. Conversely, language model head errors are non-uniform, favoring deterministic RN.
Experiments showed that pure RN increased perplexity to 2.21x the baseline, while pure SR resulted in 1.15x. By assigning SR to the MLP and RN to the head, perplexity was reduced to 1.10x, a 28% improvement over matched-bit RN.
These results are based on simulations with the small DistilGPT-2 model and have not yet been verified on large-scale models or real hardware.