Open Models Judge Math Proofs Cheaply
New research claims open-weight models like GPT-OSS match frontier LLMs in grading math proofs at a fraction of the cost.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
Open-weight models GPT-OSS-120B and DeepSeek-V4-Flash perform statistically no worse than frontier models like Claude Opus 4.7 in automated mathematical proof grading, while costing 4 to 100 times less.
Traditionally, evaluating AI mathematical reasoning relies on expensive frontier LLMs as judges. This study found that using a consensus of three cheaper open-source models serves as an effective alternative on the IMO-GradingBench benchmark.
The experiments showed that a unanimous voting rule achieved the highest precision (0.855), while majority voting yielded the highest recall (0.912). These findings replicated on the independent ProofBench dataset.
This result comes from author-run tests in a preprint (arXiv:2608.00004v2) and has not yet been independently reproduced by third parties.