VERPO's token-level credit assignment avoids GRPO training collapse
Authors self-report VERPO beats GRPO baselines on five reasoning and tool-use tasks, with no independent reproduction.
중요도국소적증거E2 미복제작성 방식간략
VERPO decomposes teacher guidance into token-level credit assignment, avoiding the optimization collapse caused by GRPO's sequence-level advantage estimation, and self-reports the highest multi-task average across five scientific reasoning and tool-use tasks.
Previously GRPO estimated advantages from whole-sequence scores, and the authors say this sequence-level granularity triggers optimization collapse.
The authors self-report larger gains on smaller models; this is a first-party benchmark with no independent reproduction.
The method includes a Fisher movement-cost controller and an FEC projection; version 4 was updated on September 23, and third-party validation is still pending.