LLM reproduces economics papers at scale, with discrepancies flagged in nearly 80%
Economists used an LLM to reproduce economics research in bulk, flagging discrepancies in 3,460 of 4,452 replication packages; flagged results are automated detections, not proof of error.
重要度重大証拠E2 未複製執筆簡易
Readers can now reproduce published economics research in bulk with a large language model: of 4,452 published replication packages, 3,460 were flagged for discrepancies, while calculations in 496 papers were sped up by more than a factor of 10 and 923 received extensions consistent with the original aims.
Previously, such reproduction relied on manual, paper-by-paper checks, which were costly and limited in coverage.
The workflow was built by economists Matthew Schwartz, Isaiah Andrews and Jesse M. Shapiro. It first reproduces the original calculations, then checks them against published findings and runs sensitivity analysis; the flagged counts and speedups are author-reported.
The flagged "discrepancies" are the workflow's automated detections, not proof the original papers were wrong; the paper has not been peer reviewed, is NBER working paper w35782, was covered by Tyler Cowen on Marginal Revolution on October 1, and has no third-party replication yet.