50,000 agent error pairs released, proposed fixes lift pass rate to 51.1%
Authors' own tests show proposed corrections raise replay verifier pass rates from 18.4% to 51.1%; failure-trace reuse is worth tracking.
重要度局所的証拠E2 未複製執筆簡易
Agent failure traces can now be reused without replaying rollouts: the Agent Error Dataset contains 50,228 error-diagnosis pairs, and in the authors' own tests proposed corrections raised replay verifier pass rates from 18.4% to 51.1%.
Previously, diagnosing failures required replaying the original tasks, which was costly and hard to reuse; the dataset retains source traces and execution metadata so failures can be re-diagnosed without replaying.
Kunlun Zhu and five co-authors released the dataset on arXiv on September 30; the pairs come from 9,961 tasks across 33 environments and 23 policy models, and the authors' own tests on 3,062 matched replay pairs produced the pass-rate result, while fine-tuning Qwen3-8B lifted exact-step agreement with teacher labels from 47.2% to 63.6%.
The replay comparison covers only environments that support replay, the results are the authors' own, and no third party has yet reproduced them.