BudgetAPO Preprint: Noise-Adaptive Evaluation Cuts Prompt Optimization Failure Rate
New algorithm dynamically adjusts evaluation slice size to significantly reduce prompt optimization failures under limited API call budgets.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Key Finding: BudgetAPO reduces the failure rate of returning the original seed prompt from 86% (for GEPA) to 13% within a 250-call budget, addressing the practical usability of Automatic Prompt Optimization (APO) under paid API constraints.
Context: Existing APO methods like GEPA and OPRO typically assume hundreds to thousands of model calls, which is impractical in rate-limited or paid API environments. Under tight budgets, multi-stage pipelines often exhaust the budget and revert to the initial state, while single-stage methods perform poorly due to fixed-size evaluation batches that ignore task-specific noise differences.
Result: The preprint proposes a noise-adaptive rule that sizes the evaluation slice based on measured task noise, using paired comparisons for accept/reject decisions. Authors' self-tests across seven benchmarks and five models show BudgetAPO ranks first on all subjects and outperforms all baselines under Holm-corrected paired tests. For instance, on GPT-OSS-20B, GEPA requires 5 times as many calls to match BudgetAPO's score achieved in 100 calls.
Boundary: Results are author-reported and not yet independently reproduced. The paper is an arXiv preprint v2 (updated Oct 8, 2026), with no indication of peer review.