Vals AI Test: Multi-Agent Teams Cost Up to 5x More with Minimal Quality Gain
Vals AI benchmarks show that while multi-agent teams can parallelize work, they cost up to 5.1x more than single agents with often insignificant quality improvements.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
AI agent teams building web applications cost 1.8 to 5.1 times more than single agents, yet deliver barely measurable quality gains.
Vals AI tested GPT-6 Sol and Claude Opus 5.5 on the Vibe Code Bench. Only Sol's medium-effort team showed a statistically significant score improvement (+7.3 points); the other three comparisons yielded no significant difference.
Higher reasoning effort for a single agent often proved more cost-effective than lower-effort teams. For instance, Sol's max-effort single agent cost 1.6x more than its medium-effort team but scored higher. Opus's team mode led to longer runtimes and skyrocketing costs without significant quality leaps.
These findings come from Vals AI's self-reported benchmark on full-stack web apps and may not generalize to other task types.