Stochastic Security: Adversarial Defense Using Long-Run Dynamics of Energy-Based Models
Mitch Hill, Jonathan Craig Mitchell, Song-Chun Zhu
Abstract
The vulnerability of deep networks to adversarial attacks is a central problem for deep learning from the perspective of both cognition and security. The current most successful defense method is to train a classifier using adversarial images created during learning. Another defense approach involves transformation or purification of the original input to remove adversarial signals before the image is classified. We focus on defending naturally-trained classifiers using Markov Chain Monte Carlo (MCMC) sampling with an Energy-Based Model (EBM) for adversarial purification. In contrast to adversarial training, our approach is intended to secure highly vulnerable pre-existing classifiers. To our knowledge, no prior defensive transformation is capable of securing naturally-trained classifiers, and our method is the first to validate a post-training defense approach that is distinct from current successful defenses which modify classifier training. The memoryless behavior of long-run MCMC sampling will eventually remove adversarial signals, while metastable behavior preserves consistent appearance of MCMC samples after many steps to allow accurate long-run prediction. Balancing these factors can lead to effective purification and robust classification. We evaluate adversarial defense with an EBM using the strongest known attacks against purification. Our contributions are 1) an improved method for training EBM's with realistic long-run MCMC samples for effective purification, 2) an Expectation-Over-Transformation (EOT) defense that resolves ambiguities for evaluating stochastic defenses and from which the EOT attack naturally follows, and 3) state-of-the-art adversarial defense for naturally-trained classifiers and competitive defense compared to adversarial training on CIFAR-10, SVHN, and CIFAR-100. Our code and pre-trained models are available at https://github.com/point0bar1/ebm-defense .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbb9a05b-89e5-4047-b59a-8c41dbac7fdfCited by top-tier papers33
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
- Adversarial Purification with Score-based Generative ModelsJongmin Yoon, Sung Ju Hwang, Juho LeeICML 2021 · 200 citations
- Oops I Took A Gradient: Scalable Sampling for Discrete DistributionsWill Grathwohl, Kevin Swersky, Milad Hashemi, David Duvenaud et al.ICML 2021 · 113 citations
- Robust Evaluation of Diffusion-Based Adversarial PurificationMinjong Lee, Dongwoo KimICCV 2023 · 96 citations
- Evaluating the Adversarial Robustness of Adaptive Test-time DefensesFrancesco Croce, Sven Gowal, Thomas Brunner, Evan Shelhamer et al.ICML 2022 · 85 citations
Builds on7
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- Denoised Smoothing: A Provable Defense for Pretrained ClassifiersHadi Salman, Mingjie Sun, Greg Yang, Ashish Kapoor et al.NeurIPS 2020 · 191 citations
Related papers
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
- Adversarial Training on Purification (AToP): Advancing Both Robustness and GeneralizationGuang Lin, Chao Li, Jianhai Zhang, Toshihisa Tanaka et al.ICLR 2024 · 25 citations
- Adversary Aware Optimization for Robust DefenseDaniel Wesego, Pedram RooshenasNeurIPS 2025 · 3 citations
- PureGen: Universal Data Purification for Train-Time Poison Defense via Generative Model DynamicsOmead Pooladzandi, Sunay Bhat, Jeffrey Jiang, Alexander Branch et al.NeurIPS 2024 · 6 citations
- Adversarial Attacks are Reversible with Natural SupervisionChengzhi Mao, Mia Chiquier, Hao Wang, Junfeng Yang et al.ICCV 2021 · 66 citations
