Agent training scored by verified progress, authors report 4.1-point gain
Progress can be verified like outcomes, so failed attempts still yield training signal; the numbers are the authors' own tests, pending third-party replication.
ImportanciaLocalEvidenciaE2 no replicadoAnálisisRápido
ProCredit lets every turn of agent training be scored by verified progress: the acceptance checks that decide task success are rerun on the intermediate state after each turn, rather than assigning a single outcome reward at the end, and the authors report a 4.1-percentage-point gain over the strongest outcome-reward baseline at 4B on AppWorld.
The known weakness of outcome rewards is that a group of attempts with no success yields no training signal, and failures cannot be told apart by how close they came.
The authors report that on the AppWorld benchmark, starting from Qwen3.5 base models at three scales, the method beats both outcome-reward and progress-based baselines at every scale; a second environment shows the same direction. Ablations indicate that adding final progress to the trajectory score alone does not help — the gain comes from crediting progress to the turn where it occurs. The numbers are the authors' own tests, the abstract gives no figures for the second environment, and independent replication is still pending.