Adversarial Purification with Score-based Generative Models
Jongmin Yoon, Sung Ju Hwang, Juho Lee
Abstract
While adversarial training is considered as a standard defense method against adversarial attacks for image classifiers, adversarial purification, which purifies attacked images into clean images with a standalone purification model, has shown promises as an alternative defense method. Recently, an Energy-Based Model (EBM) trained with Markov-Chain Monte-Carlo (MCMC) has been highlighted as a purification model, where an attacked image is purified by running a long Markov-chain using the gradients of the EBM. Yet, the practicality of the adversarial purification using an EBM remains questionable because the number of MCMC steps required for such purification is too large. In this paper, we propose a novel adversarial purification method based on an EBM trained with Denoising Score-Matching (DSM). We show that an EBM trained with DSM can quickly purify attacked images within a few steps. We further introduce a simple yet effective randomized purification scheme that injects random noises into images before purification. This process screens the adversarial perturbations imposed on images by the random noises and brings the images to the regime where the EBM can denoise well. We show that our purification method is robust against various attacks and demonstrate its state-of-the-art performances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90a56c7d-58e4-49a8-954a-18f91831d3d0Cited by top-tier papers62
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
- The Stable Signature: Rooting Watermarks in Latent Diffusion ModelsPierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze et al.ICCV 2023 · 370 citations
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin et al.ICML 2023 · 300 citations
- GENIE: Higher-Order Denoising Diffusion SolversTim Dockhorn, Arash Vahdat, Karsten KreisNeurIPS 2022 · 161 citations
- Diffusion-Based Adversarial Sample Generation for Improved Stealthiness and ControllabilityHaotian Xue, Alexandre Araujo, Bin Hu, Yongxin ChenNeurIPS 2023 · 110 citations
Builds on11
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 1,527 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
Related papers
- Stochastic Security: Adversarial Defense Using Long-Run Dynamics of Energy-Based ModelsMitch Hill, Jonathan Craig Mitchell, Song-Chun ZhuICLR 2021 · 93 citations
- PureGen: Universal Data Purification for Train-Time Poison Defense via Generative Model DynamicsOmead Pooladzandi, Sunay Bhat, Jeffrey Jiang, Alexander Branch et al.NeurIPS 2024 · 6 citations
- Text Adversarial Purification as Defense against Adversarial AttacksLinyang Li, Demin Song, Xipeng QiuACL 2023 · 11 citations
- Adversary Aware Optimization for Robust DefenseDaniel Wesego, Pedram RooshenasNeurIPS 2025 · 3 citations
- ADBM: Adversarial Diffusion Bridge Model for Reliable Adversarial PurificationXiao Li, Wenxuan Sun, Huanran Chen, Qiongxiu Li et al.ICLR 2025
