1.8M human-agent co-authored code edits self-reported to fine-tune better than human commit data
Northeastern researchers self-report: ~1.8M code edits co-authored with Claude Code, Codex and Cursor, scraped from GitHub, fine-tune better than human commit corpora on most benchmarks.
ImportanceLocalEvidenceE2 unreplicated
Readers can now download and use a corpus of roughly 1.8 million code edits (371GB) co-authored by Claude Code, OpenAI Codex, and Cursor Agent with humans, whose natural language descriptions are nearly 10x longer than prior human commit datasets.
Previously, such agent-human collaborative edit data was lacking, and training relied on purely human commit corpora, which the authors say limited fine-tuning results.
A Northeastern University team self-reports: the corpus was collected from public GitHub records from early April to October 2025, and models fine-tuned on it outperform models trained on purely human commit corpora on most of the HumanEvalFix, CanItEdit, and similar benchmarks; the corpus is released on HuggingFace (nuprl/AgentPack).
Results are author self-reported, with no third-party reproduction yet; the work first appeared September 26, 2025 and was updated to v3 on September 16, 2026 (arXiv:2509.21891).