Small-model experiment says a few dozen samples can silently rewrite model behavior
Paper under review, model only 20M parameters; whether it generalizes to real large models remains unverified.
The author's own experiment reports that after a few hundred samples teach the model to bypass a shortcut, roughly 56 wrong-answer samples flip half of the questions to the longer route, while accuracy on normal tests stays perfect.
The experiment ran on a 20-million-parameter model with a manually constructed verb-question task. Scaling the training set from 32,000 to 320,000 samples left the required special samples at a few hundred (about 378 at 96,000), meaning the absolute count matters, not the proportion. This matches Souly et al.'s 2025 finding that pre-training poisoning takes about 250 documents regardless of model or data size.
In the September 30 LessWrong post the author offers an explanation: ordinary samples have a capped strengthening effect on the long route, and only samples where the shortcut is blocked or wrong truly reinforce it. The paper is under review; whether the finding generalizes to real large models remains unverified.