Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
Shengwei An, Sheng-Yen Chou, Kaiyuan Zhang, Qiuling Xu, Guanhong Tao, Guangyu Shen, Siyuan Cheng, Shiqing Ma, Pin-Yu Chen, Tsung-Yi Ho, Xiangyu Zhang
Abstract
Diffusion models (DM) have become state-of-the-art generative models because of their capability of generating high-quality images from noises without adversarial training. However, they are vulnerable to backdoor attacks as reported by recent studies. When a data input (e.g., some Gaussian noise) is stamped with a trigger (e.g., a white patch), the backdoored model always generates the target image (e.g., an improper photo). However, effective defense strategies to mitigate backdoors from DMs are underexplored. To bridge this gap, we propose the first backdoor detection and removal framework for DMs. We evaluate our framework ELI-JAH on over hundreds of DMs of 3 types including DDPM, NCSN and LDM, with 13 samplers against 3 existing backdoor attacks. Extensive experiments show that our approach can have close to 100% detection accuracy and reduce the backdoor effects to close to zero without significantly sacrificing the model utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce31143c-8a35-4623-87fb-cfa7cc098576Cited by top-tier papers7
- Label-Free Backdoor Attacks in Vertical Federated LearningWei Shen, Wenke Huang, Guancheng Wan, Mang YeAAAI 2025 · 15 citations
- Exploring the Orthogonality and Linearity of Backdoor AttacksKaiyuan Zhang, Siyuan Cheng, Guangyu Shen, Guanhong Tao et al.S&P 2024 · 11 citations
- TWIST: Text-encoder Weight-editing for Inserting Secret Trojans in Text-to-Image ModelsXindi Li, Zhe Liu, Tong Zhang, Jiahao Chen et al.ACL 2025 · 1 citation
- Watch the Watchers! On the Security Risks of Robustness-Enhancing Diffusion ModelsChangjiang Li, Ren Pang, Bochuan Cao, Jinghui Chen et al.USENIX Security 2025
- Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion ModelsSangwon Jang, June Suk Choi, Jaehyeong Jo, Kimin Lee et al.CVPR 2025
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- TERD: A Unified Framework for Safeguarding Diffusion Models Against BackdoorsYichuan Mo, Hui Huang, Mingjie Li, Ang Li et al.ICML 2024 · 31 citations
- UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion ModelsYuning Han, Bingyin Zhao, Rui Chu, Feng Luo et al.CVPR 2025
- VillanDiffusion: A Unified Backdoor Attack Framework for Diffusion ModelsSheng-Yen Chou, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 101 citations
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- How to Backdoor Diffusion Models?Sheng-Yen Chou, Pin-Yu Chen, Tsung-Yi HoCVPR 2023
