Leipzig Math Benchmark: Only 1 Unsolved
Next-gen LLMs solved 99 of 100 research-level math questions, nearing expert capability.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Next-generation large language models have left only one of 100 research-level mathematics questions unsolved in the "Leipzig Benchmark."
The dataset was compiled by 49 mathematicians between April and May 2026 to test AI capabilities on high-difficulty problems with known answers. Earlier evaluation stages showed 41 questions completely unsolved in initial attempts, dropping to 2 after multi-run evaluations and heavy-thinking model interventions.
This update (Stage 4) introduced configurations of next-generation models equipped with web search and code execution. The results indicate that the vast majority of questions are now conquered, suggesting that current frontier models approach human-expert levels in specific structured mathematical reasoning tasks.
Note that these are author-reported results on a small sample size (100 questions), and independent third-party reproduction has not yet been observed.