TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
Yichuan Mo, Hui Huang, Mingjie Li, Ang Li, Yisen Wang
Abstract
Diffusion models have achieved notable success in image generation, but they remain highly vulnerable to backdoor attacks, which compromise their integrity by producing specific undesirable outputs when presented with a pre-defined trigger. In this paper, we investigate how to protect diffusion models from this dangerous threat. Specifically, we propose TERD, a backdoor defense framework that builds unified modeling for current attacks, which enables us to derive an accessible reversed loss. A trigger reversion strategy is further employed: an initial approximation of the trigger through noise sampled from a prior distribution, followed by refinement through differential multi-step samplers. Additionally, with the reversed trigger, we propose backdoor detection from the noise space, introducing the first backdoor input detection approach for diffusion models and a novel model detection algorithm that calculates the KL divergence between reversed and benign distributions. Extensive evaluations demonstrate that TERD secures a 100% True Positive Rate (TPR) and True Negative Rate (TNR) across datasets of varying resolutions. TERD also demonstrates nice adaptability to other Stochastic Differential Equation (SDE)-based models. Our code is available at https://github.com/PKU-ML/TERD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b477cacc-366a-4da6-bc5f-27973b32a623Cited by top-tier papers10
- CopyrightShield: Enhancing Diffusion Model Security Against Copyright Infringement AttacksZhixiang Guo, Siyuan Liang, Aishan Liu, Dacheng TaoICCV 2025 · 8 citations
- Customization under Fire: Plugin Poisoning in Text-to-Image EcosystemJiahao Chen, Xing He, Yong Yang, Xinfeng Li et al.CCS 2026 · 2 citations
- BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response DeviationFeiran Li, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.CVPR 2026 · 2 citations
- Efficient Input-Level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation VariationShengfang Zhai, Jiajun Li, Yue Liu, Huanran Chen et al.ICCV 2025 · 2 citations
- TWIST: Text-encoder Weight-editing for Inserting Secret Trojans in Text-to-Image ModelsXindi Li, Zhe Liu, Tong Zhang, Jiahao Chen et al.ACL 2025 · 1 citation
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
Related papers
- Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution ShiftShengwei An, Sheng-Yen Chou, Kaiyuan Zhang, Qiuling Xu et al.AAAI 2024 · 48 citations
- UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion ModelsYuning Han, Bingyin Zhao, Rui Chu, Feng Luo et al.CVPR 2025
- How to Backdoor Diffusion Models?Sheng-Yen Chou, Pin-Yu Chen, Tsung-Yi HoCVPR 2023
- DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent DiffusionHossein Mirzaei, Zeinab Taghavi, Sepehr Rezaee, Masoud Hadi et al.ICCV 2025 · 3 citations
- VillanDiffusion: A Unified Backdoor Attack Framework for Diffusion ModelsSheng-Yen Chou, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 101 citations
