Backdoors Bypass Image Model Erasure
New research reveals 'Erasure Evasion Backdoors' survive safety filtering, exposing banned content with up to 94% success.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Safety erasure mechanisms in text-to-image models contain a critical blind spot: models with embedded backdoors can still generate prohibited content after undergoing concept removal.
Industry consensus held that fine-tuning severs links to harmful concepts, but researchers from TU Darmstadt introduced the 'Erasure Evasion Backdoor' (EEB), showing attackers can bind triggers to target concepts so malicious links survive subsequent erasure.
In tests against six state-of-the-art erasure methods, EEB achieved 82% success against celebrity identity unlearning and 94% for object erasure, amplifying explicit content exposure by 16 times. These results come from a preprint self-test and await independent reproduction.