A few dozen samples can silently rewrite a small model's behavior
On a 20M-parameter model, about 56 wrong samples flip half the answers' route; generalization to real large models remains unverified.
ImportânciaLocalEvidênciaE2 não replicadaTratamentoRápido
About 56 wrong-answer samples can flip a 20-million-parameter model to a different answering route on half the questions while accuracy on normal tests stays perfect — the author's own reported result, provided a few hundred samples first teach the model to bypass a shortcut.
Such behavior rewriting was previously thought to require large amounts of data: Souly et al. found in 2025 that pre-training poisoning takes about 250 documents regardless of model or data size. In this experiment, scaling the training set from 32,000 to 320,000 samples left the required special samples at a few hundred (about 378 at 96,000), meaning the absolute count matters, not the proportion; the task was a manually constructed verb-question set.
In the September 30 LessWrong post the author offers an explanation: ordinary samples have a capped strengthening effect on the long route, and only samples where the shortcut is blocked or wrong truly reinforce it. The paper is under review; whether the finding generalizes to real large models remains unverified.