Preprint claims CoAx locates self-repair backup circuits; results unreplicated
Authors report 0.941 AUC recovering backup heads on GPT-2's IOI circuit and tests across 8 models, but all results are first-party and unreplicated.
Original event 2026-09-23
A preprint, arXiv 2607.01940 (v3, 23 Sep 2026), introduces conditional co-ablation (CoAx): after ablating a primary component set, rank candidates by growth in ablation effect to find backups the intervention activates.
On GPT-2-small's IOI circuit, the authors report CoAx recovers the documented backup heads at 0.941 ROC-AUC, versus 0.815 for the strongest intact-state attribution baseline; adding the selected heads to the incomplete circuit cuts incompleteness from 0.75 to 0.21. They also report it beats matched random completions on 8 non-GPT-2 models across 6 architecture families.
All numbers are first-party evaluations, backup-head recovery is mainly validated on the known-answer IOI setting, and nothing has been independently reproduced. For readers doing circuit analysis or safety evaluation, the thing to watch is whether conditional completion holds up in other groups' experiments.