AI Science Reproduction Rate Hits 14%
UniPat AI's PaperBenchX benchmark reveals that even the top-performing GPT-6 Astra achieves only a 13.98% full reproduction rate across 93 cross-disciplinary paper reproduction tasks.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upStandard
UniPat AI's newly released PaperBenchX benchmark shows that GPT-6 Astra, the best-performing model tested, achieved a full reproduction rate of only 13.98% across 93 cross-disciplinary scientific paper reproduction tasks.
The benchmark requires AI agents to rebuild and execute scientific workflows from scratch in domain-native software environments across fields like electromagnetics, chemistry, and biology, rather than merely guessing results. Verification is conducted by deleting outputs, blocking network access, and replaying the workflow to ensure evidence regenerability.
While average scores for modeling (62.70%) and execution (61.84%) stages were relatively high, validation scores dropped significantly to 42.80%, indicating that AI struggles to connect these steps into a credible end-to-end reproduction. Currently, only 12 representative tasks are open-sourced, while 81 remain restricted to maintain benchmark discrimination.