Preprint says per-token credit assignment prevents training collapse
Authors self-report wins over GRPO baselines on five reasoning and tool-use tasks; unverified, track as a method signal.
Original event 2026-09-23
A preprint introduces VERPO, which decomposes evidence-conditioned teacher guidance into token-level credit assignment; the authors say it avoids the optimization collapse caused by GRPO's sequence-level advantages.
The authors self-report the highest multi-task average across five scientific reasoning and tool-use tasks, with larger gains on smaller models. This is a first-party benchmark with no independent reproduction.
The method includes a Fisher movement-cost controller and an FEC projection; version 4 was updated on September 23, and third-party validation is still pending.