Stripping synthetic markers lets implanted false facts evade linear probes
First-party results advance the stealth of SDF-implanted false facts, but only at the probe level and without downstream follow-through.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
Reducing "synthetic markers" in synthetic document finetuning (SDF) makes some implanted false facts indistinguishable from pretraining knowledge under a linear probe. Researcher Jason Zeng published the first-party results on LessWrong on October 4.
Earlier work by Slocum et al. in 2025 showed SDF-implanted false facts are almost always caught by middle-layer linear probes, without pinning down whether the cause is synthetic-looking text or disbelief. After removing markers such as surprise framing and excess em dashes from training documents, subtly implausible false facts scored probe error rates near 0.5, meaning true and false answers became hard to tell apart.
The results are based on Llama-3.1-8B-Instruct and the author's own probe procedure; even when probes can no longer distinguish them, the false facts do not always propagate to downstream reasoning tasks, and overtraining still leaves detectable traces.