Stripping synthetic markers lets implanted false facts evade linear probes
First-party results advance the stealth of SDF-implanted false facts, but only at the probe level and without downstream follow-through.
중요도국소적증거E2 미복제작성 방식간략
Reducing "synthetic markers" in synthetic document finetuning (SDF) makes some implanted false facts indistinguishable from pretraining knowledge under a linear probe. Researcher Jason Zeng published the first-party results on LessWrong on October 4.
Earlier work by Slocum et al. in 2025 showed SDF-implanted false facts are almost always caught by middle-layer linear probes, without pinning down whether the cause is synthetic-looking text or disbelief. After removing markers such as surprise framing and excess em dashes from training documents, subtly implausible false facts scored probe error rates near 0.5, meaning true and false answers became hard to tell apart.
The results are based on Llama-3.1-8B-Instruct and the author's own probe procedure; even when probes can no longer distinguish them, the false facts do not always propagate to downstream reasoning tasks, and overtraining still leaves detectable traces.