Preprint claims agents improve success using only self-modification records
SelfSearch improves agents without reward signals, matching the top harness for about $4 in search cost, but the numbers are author-reported and unreplicated.
A preprint called SelfSearch reports that agents can improve without reward signals, using only records of past self-modification attempts, and that success rose over the initial agent in all six model-benchmark settings.
The paper by Jungwoo Yang, Injin Kong and Yohan Jo was submitted to arXiv on September 29. Instead of repeated downstream evaluation, the agent reads the reasoning, tool actions and outcomes of earlier modification attempts as its basis for improving.
Author-reported numbers include a gain of up to 11.2 percentage points on Terminal-Bench 2.1, and on SWE-bench Multilingual a 5.0-point success gain with 38.5% lower execution cost. With DeepSeek V4 Flash, $4.03 of search cost produced a harness solving 82.0% of Terminal-Bench 2.1, matching Codex, the top scorer in a public nine-harness comparison.
The abstract does not say whether the 82.0% comparison reuses the original settings of the public nine-harness comparison or the authors' own evaluation setup.
Sources:https://arxiv.org/abs/2609.37968