Masked self-distillation internalizes chain-of-thought, making Qwen3 reasoning shorter and cheaper
Authors' own tests on Qwen3-4B and Qwen3-8B report performance held or improved after internalizing chain-of-thought; results await independent replication.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
Masked self-distillation internalizes the chain of thought into model parameters: on Qwen3-4B and Qwen3-8B, the authors' own tests report task performance held or improved while the model generates fewer intermediate steps, cutting inference latency and cost.
Previously, a chain of thought is the long sequence of explicit reasoning steps a large reasoning model produces before its final answer, and these traces dominate serving latency and compute. The method instantiates copies of the same model as teacher and student, training the student to internalize part or all of the intermediate trace and emit shorter outputs.
The authors ran controlled experiments in two domains, math and graph coloring, reporting no catastrophic forgetting out of domain; an ablation found plain supervised fine-tuning also shortens traces but generalizes worse. All results are the authors' own, and specific numbers require reading the full text; the preprint was first submitted June 18 and revised to v3 on September 30, with no independent replication yet.