Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
Tobias Braun, Jonas Henry Grebe, Patrick Mohr Gordillo, Marcus Rohrbach, Anna Rohrbach
Abstract
The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery. To mitigate such risks, prior work has proposed concept erasure methods that aim to sever unwanted concepts from the model via fine-tuning, yet it remains unclear whether these approaches truly remove all links to the harmful concept or merely conceal superficial connections. In this work, we reveal a critical vulnerability, the Erasure Evasion Backdoor (EEB): an adversary binds a backdoor trigger to a concept slated for removal, and this malicious link survives subsequent erasure. We show that both black-box and white-box adversaries can instantiate this threat. Across six state-of-the-art erasure methods, including robust ones that explicitly search for alternative representations of the target concept, EEB consistently exposes harmful content: up to 82% success against celebrity-identity unlearning, up to 94% for object erasure, and up to 16 amplification of explicit-content exposure. While EEB uncovers a blind spot in current erasure methods, it also provides a diagnostic tool for stress-testing future concept erasure techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9b2412f-7a67-46e7-9771-13409a2ccdfeBuilds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Circumventing Concept Erasure Methods For Text-To-Image Generative ModelsMinh Pham, Kelly O. Marshall, Niv Cohen, Govind Mittal et al.ICLR 2024 · 82 citations
- Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation MisdirectionUichan Lee, Jeonghyeon Kim, Sangheum HwangICLR 2026 · 3 citations
- Memories of Forgotten ConceptsMatan Rusanovsky, Shimon Malnick, Amir Jevnisek, Ohad Fried et al.CVPR 2025
- TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion ModelsRuidong Chen, Honglin Guo, Lanjun Wang, Chenyu Zhang et al.ICCV 2025 · 17 citations
- Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data PoisoningShengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu et al.ACM MM 2023 · 46 citations
