VoS Boosts Agent Correction Accuracy
New research proposes VoS, using counterfactual data to learn intervention timing, improving agent performance by an average of 7.8 points.
ImportanciaMaterialEvidenciaE2 no replicadoAnálisisRápido
The Value of Steering (VoS) method improves large language model agent execution by an average of 7.8 points across 12 test settings.
Traditional approaches rely on single-step uncertainty signals to decide when to correct agents, but research shows these signals cannot reliably locate effective intervention steps. VoS analyzes approximately 82,000 counterfactual continuations to learn the value of steering at each step, introducing a harm budget to limit interference with successful trajectories.
The method outperforms the strongest of five existing uncertainty-triggered strategies by an average of 2.9 points. Results are from a preprint's self-testing and have not yet been independently reproduced.