AI produces 36 science manuscripts in three months, expert review still essential
A Harvard team used an open-source harness to run Claude across 18 fields; output was large but every result needed expert verification.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Harvard physicist Matthew Schwartz used the open-source BootLoops harness to drive Claude, producing 36 manuscripts with 19 co-authors in three months, spanning 18 fields from particle physics to linguistics.
Concrete results include 30 elliptic integrals computed by Claude, fifteen of them for the first time, and a solution to a 20-year-old equation in neutral biodiversity theory that had been impossible to compute at scale; applied to data, it showed tree species composition on Barro Colorado Island in Panama changing 4.5 times faster than the theory allows.
In his guest post on Anthropic's site, Schwartz warns the model declares victory early, automated checks are unreliable, and conclusions can be wrong even when calculations are correct — each result became scientifically valuable only after domain experts stepped in. The manuscripts have not been peer reviewed.