Preprint claims self-distillation training recovers losses from extreme visual compression
Authors' own tests show performance rising from 68.6% to 82.3% at 5% visual tokens; worth tracking, pending independent replication.
A preprint called LT-OPD proposes a self-distillation training method: the compressed multimodal model learns from trajectories it generates itself, with a full-visual-token copy of the same model serving as teacher, while the token budget is progressively lowered during training.
In the authors' own tests on Qwen3.5-4B, average retained performance across nine benchmarks rose from 68.6% to 82.3% at 5% visual-token retention, while KV-cache usage fell 85.2% and prefill FLOPs fell 85.4%. The paper reports consistent gains on Qwen3.5-9B, GLM-4.6V-9B and LLaVA-OV-1.5-4B.
The paper was submitted to arXiv on September 26; all results are the authors' own tests on self-chosen benchmarks and await independent replication.
Sources:https://arxiv.org/abs/2609.32353