Nvidia says tuning the harness, not the model, can nearly halve coding-agent token use
Automated harness optimization moves agent cost-cutting off the model, but gains concentrate on the authors' chosen benchmark.
重要度局所的証拠E2 未複製執筆簡易
A coding agent's harness, once automatically optimized, can cut token use by 49 percent in its leanest variant while retaining 93.7 percent of the Pi harness's score — the result reported for SoL-Pi, a system built by Nvidia researchers.
Cost-cutting previously focused on the model side, with harnesses rarely rewritten systematically by automation. SoL-Pi uses a research agent to explore 152 directions across 535 executable environments and more than 3,000 runs, producing four mechanisms including action fusion and context compaction. At current API prices, the authors estimate savings of $8.75 to $13.50 per hour versus native Codex and Claude Code harnesses.
The gains do not hold everywhere: on 63 Terminal-Bench 4 tasks, SoL-Pi solved 15 while Codex and Pi each solved 18. The authors themselves note that automatically optimized harnesses tend to overfit their training tasks. The result comes from a preprint submitted on 2026-09-26 and has not yet been independently reproduced.