Code Agent Eval Flaw: 'Lucky Passes' Found
New study reveals 10.7% of passing code agent trajectories have flawed processes, potentially shifting model rankings by up to five positions.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Current evaluation of software engineering (SWE) agents relies solely on whether the final patch passes tests, a standard now shown to be blind.
The AgentLens team analyzed 2,614 OpenHands trajectories and found that 10.7% of passing cases were "Lucky Passes." These trajectories exhibited chaotic behaviors such as regression cycles, blind retries, or missing verification, indicating unreliable processes despite correct outcomes.
When ranked by process quality instead of pass rate, some models shifted by as many as five rank positions. The study argues that binary signals cannot distinguish principled solutions from trial-and-error luck, advocating for process-level assessment frameworks.