Adversarial Attacks are Reversible with Natural Supervision
Chengzhi Mao, Mia Chiquier, Hao Wang, Junfeng Yang, Carl Vondrick
Abstract
We find that images contain intrinsic structure that enables the reversal of many adversarial attacks. Attack vectors cause not only image classifiers to fail, but also collaterally disrupt incidental structure in the image. We demonstrate that modifying the attacked image to restore the natural structure will reverse many types of attacks, providing a defense. Experiments demonstrate significantly improved robustness for several state-of-the-art models across the CIFAR-10, CIFAR-100, SVHN, and ImageNet datasets. Our results show that our defense is still effective even if the attacker is aware of the defense mechanism. Since our defense is deployed during inference instead of training, it is compatible with pre-trained networks as well as most other defenses. Our results suggest deep networks are vulnerable to adversarial examples partly because their representations do not enforce the natural structure of images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e10d7222-4673-49f2-b51e-9ffe0975e0d0Cited by top-tier papers22
- MEMO: Test Time Robustness via Adaptation and AugmentationMarvin Zhang, Sergey Levine, Chelsea FinnNeurIPS 2022 · 595 citations
- Evaluating the Adversarial Robustness of Adaptive Test-time DefensesFrancesco Croce, Sven Gowal, Thomas Brunner, Evan Shelhamer et al.ICML 2022 · 85 citations
- DISCO: Adversarial Defense with Local Implicit FunctionsChih-Hui Ho, Nuno VasconcelosNeurIPS 2022 · 65 citations
- Bayesian Invariant Risk MinimizationYong Lin, Hanze Dong, Hao Wang, Tong ZhangCVPR 2022 · 48 citations
- Self-Interpretable Time Series Prediction with Counterfactual ExplanationsJingquan Yan, Hao WangICML 2023 · 29 citations
Builds on20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
Related papers
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
- Defending Against Universal Attacks Through Selective Feature RegenerationTejas S. Borkar, Felix Heide, Lina J. KaramCVPR 2020
- Stochastic Security: Adversarial Defense Using Long-Run Dynamics of Energy-Based ModelsMitch Hill, Jonathan Craig Mitchell, Song-Chun ZhuICLR 2021 · 93 citations
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 23 citations
- DIPDefend: Deep Image Prior Driven Defense against Adversarial ExamplesTao Dai, Yan Feng, Dongxian Wu, Bin Chen et al.ACM MM 2020 · 20 citations
