Independent audit finds weight-decomposition explanations drift when aggregated
An independent audit of Goodfire's VPD decomposition finds aggregation and deletion operations measurably shift model outputs; whether this holds at production scale is untested.
An independent audit finds Goodfire's weight-decomposition explanations drift clearly from the original model once aggregated.
Adversarial Parameter Decomposition (VPD) splits model weights into simple components and labels each as "needed here" or "safe to remove" per token. On September 30, Tom Angsten published an audit on LessWrong: after aggregating components for 64 tokens (about 4,500), output drift reached 0.80 nats, close to the 0.83 caused by the paper's own 20-step adversary.
Deleting every component never labeled as needed moved the model 1.28 nats in KL. The audit covers only the decomposition released with the published paper, not the newer unpublished training recipe in the authors' repository; the subject is a four-layer, 67M-parameter model, and whether this holds at production scale is unknown.
Sources:https://www.lesswrong.com/posts/KeBccWBGXnNXzZFBp/do-vpd-s-explanations-aggregate-an-audit-of-the-released-1https://github.com/angsten/vpd-audit