New self-distillation method modestly improves computer-use agent training
The method turns execution feedback into step-level training signals; self-reported gains are small and await independent replication.
ComputerSD converts real-time feedback from executed actions into training signals for computer-use agents; its authors report it beats training with outcome rewards only on the OSWorld-Verified benchmark.
The numbers: 1.9 percentage points higher on the Qwen3-VL-8B-Thinking backbone and 4.1 points higher on the specialized EvoCUA-8B, all measured by the paper's own authors.
The gains are modest and the benchmark was run by the authors themselves; whether the method delivers comparable benefits in real use awaits independent replication.
Sources:https://arxiv.org/abs/2609.40253https://github.com/ZJU-REAL/ComputerSD