50,000 agent error pairs released, proposed fixes lift pass rate to 51.1%
Authors' own tests show proposed corrections raise replay verifier pass rates from 18.4% to 51.1%; failure-trace reuse is worth tracking.
Kunlun Zhu and five co-authors released the Agent Error Dataset on arXiv on September 30, containing 50,228 error-diagnosis pairs.
The pairs come from 9,961 tasks across 33 environments and 23 policy models, with source traces and execution metadata retained so failures can be re-diagnosed without replaying rollouts.
In the authors' own tests, proposed corrections raised verifier pass rates from 18.4% to 51.1% across 3,062 matched replay pairs, and fine-tuning Qwen3-8B lifted exact-step agreement with teacher labels from 47.2% to 63.6%. The replay comparison covers only environments that support replay.
Sources:https://arxiv.org/abs/2609.40111