Data Poisoning Survives Filtering
New research shows 'Phantom Transfer' attacks evade 11 data-level defenses, even when the poisoning method is known.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
The 'Phantom Transfer' data poisoning attack evades 11 tested data-level defenses, including paraphrasing every sample.
The attack uses modified subliminal learning to embed malicious behaviors in training data. Even if developers know exactly how the poison was placed into an otherwise benign dataset, current filtering mechanisms cannot remove it.
Experiments demonstrate effectiveness regardless of the source model, target model, or attack goal. Authors suggest future defenses must supplement data cleaning with white-box methods and post-training audits.
This is a self-reported result from a NeurIPS 2026 accepted paper; the abstract does not specify whether tests were simulated or on real hardware.