Preprint claims training-free pruning can cut half of MoE experts while leading on reasoning scores
RAZOR scores experts by replaceability rather than contribution; the authors' own tests lead, and reproduction will decide its deployment value.
A preprint called RAZOR proposes a training-free pruning method for mixture-of-experts models; the authors' own tests show it leading comparable methods across four models and eight settings.
MoE models activate only a few experts per token yet store the entire pool. Authors Mingyang Song and Mao Zheng argue that deletion damage is decided not by an expert's contribution size but by whether surviving computation can reproduce its output; RAZOR computes the exact output change from deleting one expert using "consensus residuals", forward passes only, with no gradients or recovery training.
Removing 25% and 50% of experts on GLM-4.7-Flash, Qwen3.6-35B-A3B and two other models, the authors report RAZOR leads on the nine-task reasoning average in all settings, beating REAP by 2.12 to 5.59 points. The paper itself notes pruned models still shift in response diversity, formatting and termination.
The abstract does not list the full set of pruning methods compared, so the lead holds only against the few baselines evaluated in the paper.
Sources:https://arxiv.org/abs/2609.30465