Self-distillation training recovers 13.7 points at 5% visual tokens
LT-OPD's authors' own tests show Qwen3.5-4B's average retained performance rising from 68.6% to 82.3% at 5% visual tokens, pending independent replication.
ImportânciaLocalEvidênciaE2 não replicadaTratamentoRápido
A self-distillation training method substantially recovers the losses of extreme visual compression: on Qwen3.5-4B with only 5% of visual tokens retained, average retained performance across nine benchmarks rose from 68.6% to 82.3%, while KV-cache usage fell 85.2% and prefill FLOPs fell 85.4%.
Previously, compressed multimodal models suffered severe losses at such low token budgets, with no training approach that could restore performance while compressing.
The method, called LT-OPD, has the compressed model learn from trajectories it generates itself, with a full-visual-token copy of the same model serving as teacher, while the token budget is progressively lowered during training; the authors' own tests report consistent gains on Qwen3.5-9B, GLM-4.6V-9B and LLaVA-OV-1.5-4B.
All results are the authors' own tests on self-chosen benchmarks and await independent replication; the method was submitted to arXiv as a preprint on September 26.