Study finds repairing an agent's failed experience is not the same as reusing it
A preprint experiment suggests agent memory updates need both a previous-version reference and a fresh-start reference; findings await independent replication.
A preprint experiment finds that the value of repairing an agent's failed experience is distinct from the value of reusing it.
The paper by Yanfei Zhang and Xu Lin was submitted to arXiv on September 28, covering 3,300 runs under eleven conditions. ThinkingBox's Full mode shows a 44-percentage-point correction gain, yet 12 of its 15-point edge over Skill comes from worse uncorrected performance, not better corrected memory.
The authors conclude memory updates require both a previous-version reference and a fresh-start reference. All experiments ran on the authors' own ThinkingBox and APEX frameworks, with no third-party replication.
Sources:https://arxiv.org/abs/2609.34603