Diffusion Models for Adversarial Purification
Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, Animashree Anandkumar
Abstract
Adversarial purification refers to a class of defense methods that remove adversarial perturbations using a generative model. These methods do not make assumptions on the form of attack and the classification model, and thus can defend pre-existing classifiers against unseen threats. However, their performance currently falls behind adversarial training methods. In this work, we propose DiffPure that uses diffusion models for adversarial purification: Given an adversarial example, we first diffuse it with a small amount of noise following a forward diffusion process, and then recover the clean image through a reverse generative process. To evaluate our method against strong adaptive attacks in an efficient and scalable way, we propose to use the adjoint method to compute full gradients of the reverse generative process. Extensive experiments on three image datasets including CIFAR-10, ImageNet and CelebA-HQ with three classifier architectures including ResNet, WideResNet and ViT demonstrate that our method achieves the state-of-the-art results, outperforming current adversarial training and adversarial purification methods, often by a large margin. Project page: https://diffpure.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a83829a-95d0-4c53-bc79-bb6b80662235Cited by top-tier papers225
- PhysDiff: Physics-Guided Human Motion Diffusion ModelYe Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat et al.ICCV 2023 · 414 citations
- Diffusion Models as Plug-and-Play PriorsAlexandros Graikos, Nikolay Malkin, Nebojsa Jojic, Dimitris SamarasNeurIPS 2022 · 323 citations
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin et al.ICML 2023 · 300 citations
- I2SB: Image-to-Image Schrödinger BridgeGuan-Horng Liu, Arash Vahdat, De-An Huang, Evangelos A. Theodorou et al.ICML 2023 · 252 citations
- Invisible Image Watermarks Are Provably Removable Using Generative AIXuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan et al.NeurIPS 2024 · 209 citations
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
Related papers
- Diffusion Models Demand Contrastive Guidance for Adversarial Purification to AdvanceMingyuan Bai, Wei Huang, Tenghui Li, Andong Wang et al.ICML 2024 · 18 citations
- ADBM: Adversarial Diffusion Bridge Model for Reliable Adversarial PurificationXiao Li, Wenxuan Sun, Huanran Chen, Qiongxiu Li et al.ICLR 2025
- MimicDiffusion: Purifying Adversarial Perturbation via Mimicking Clean Diffusion ModelKaiyu Song, Hanjiang Lai, Yan Pan, Jian YinCVPR 2024 · 11 citations
- Defending against Adversarial Audio via Diffusion ModelShutong Wu, Jiongxiao Wang, Wei Ping, Weili Nie et al.ICLR 2023 · 6 citations
- Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial PurificationGaozheng Pei, Shaojie Lyu, Gong Chen, Ke Ma et al.CVPR 2025
