Repairing an agent's failed experience is not the same as reusing it
Experiments show 12 of ThinkingBox Full mode's 15-point edge over Skill comes from worse uncorrected performance, within a 44-point correction gain; findings await independent replication.
ImportanceLocalPreuvesE2 non réplicableTraitementRapide
The value of repairing an agent's failed experience is distinct from the value of reusing it: ThinkingBox's Full mode shows a 44-percentage-point correction gain, yet 12 of its 15-point edge over Skill comes from worse uncorrected performance, not better corrected memory.
Looking only at correction gains can mistake repair value for reuse value, overstating the effect of memory updates.
Experiments by Yanfei Zhang and Xu Lin, submitted to arXiv on September 28, covered 3,300 runs under eleven conditions; the authors conclude memory updates require both a previous-version reference and a fresh-start reference.
All experiments ran on the authors' own ThinkingBox and APEX frameworks, with no third-party replication.