Agents self-improve without reward signals, matching top harness for $4
SelfSearch lifts agent success using only past self-modification records, matching top public harness Codex for $4.03 in search cost, but the numbers are author-reported and unreplicated.
ImportânciaLocalEvidênciaE2 não replicadaTratamentoRápido
You can now improve agents without reward signals, using only records of past self-modification attempts: SelfSearch raised success over the initial agent in all six model-benchmark settings, and with DeepSeek V4 Flash, $4.03 of search cost produced a harness solving 82.0% of Terminal-Bench 2.1.
Improving agents previously meant repeated downstream evaluation, which is costly. SelfSearch has the agent read the reasoning, tool actions and outcomes of earlier modification attempts as its basis for improving, removing that cost.
Author-reported numbers include a gain of up to 11.2 percentage points on Terminal-Bench 2.1, and on SWE-bench Multilingual a 5.0-point success gain with 38.5% lower execution cost. The $4.03 harness matches Codex, the top scorer in a public nine-harness comparison.
The abstract does not say whether the 82.0% comparison reuses the original settings of the public nine-harness comparison or the authors' own evaluation setup; the paper by Jungwoo Yang, Injin Kong and Yohan Jo was submitted to arXiv on September 29, is author-reported and awaits replication.