One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial Examples
Chang Xiao, Changxi Zheng
Abstract
Modern image classification systems are often built on deep neural networks, which suffer from adversarial examples-images with deliberately crafted, imperceptible noise to mislead the network's classification. To defend against adversarial examples, a plausible idea is to obfuscate the network's gradient with respect to the input image. This general idea has inspired a long line of defense methods. Yet, almost all of them have proven vulnerable. We revisit this seemingly flawed idea from a radically different perspective. We embrace the omnipresence of adversarial examples and the numerical procedure of crafting them, and turn this harmful attacking process into a useful defense mechanism. Our defense method is conceptually simple: before feeding an input image for classification, transform it by finding an adversarial example on a pretrained external model. We evaluate our method against a wide range of possible attacks. On both CIFAR-10 and Tiny ImageNet datasets, our method is significantly more robust than state-of-the-art methods. Particularly, in comparison to adversarial training, our method offers lower training cost as well as stronger robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8cf5799-049f-4dc4-b00a-86e7f70465aeCited by top-tier papers2
- Stereoscopic Universal Perturbations across Different Architectures and DatasetsZachary Berger, Parth Agrawal, Tian Yu Liu, Stefano Soatto et al.CVPR 2022 · 8 citations
- Delving into Deep Image Prior for Adversarial Defense: A Novel Reconstruction-based Defense FrameworkLi Ding, Yongwei Wang, Xin Ding, Kaiwen Yuan et al.ACM MM 2021 · 7 citations
Builds on5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Adversarial Defense by Restricting the Hidden Space of Deep Neural NetworksAamir Mustafa, Salman H. Khan, Munawar Hayat, Roland Goecke et al.ICCV 2019 · 160 citations
- Bilateral Adversarial Training: Towards Fast Training of More Robust Models Against Adversarial AttacksJianyu Wang, Haichao ZhangICCV 2019 · 120 citations
- Enhancing Adversarial Defense by k-Winners-Take-AllChang Xiao, Peilin Zhong, Changxi ZhengICLR 2020 · 114 citations
Related papers
- Adversarial Attacks are Reversible with Natural SupervisionChengzhi Mao, Mia Chiquier, Hao Wang, Junfeng Yang et al.ICCV 2021 · 66 citations
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 23 citations
- Defending Against Universal Attacks Through Selective Feature RegenerationTejas S. Borkar, Felix Heide, Lina J. KaramCVPR 2020
- DIPDefend: Deep Image Prior Driven Defense against Adversarial ExamplesTao Dai, Yan Feng, Dongxian Wu, Bin Chen et al.ACM MM 2020 · 20 citations
- Fake Gradient: A Security and Privacy Protection Framework for DNN-based Image ClassificationXianglong Feng, Yi Xie, Mengmei Ye, Zhongze Tang et al.ACM MM 2021 · 1 citation
