GAT: Generative Adversarial Training for Adversarial Example Detection and Robust Classification
Xuwang Yin, Soheil Kolouri, Gustavo K. Rohde
Abstract
The vulnerabilities of deep neural networks against adversarial examples have become a significant concern for deploying these models in sensitive domains. Devising a definitive defense against such attacks is proven to be challenging, and the methods relying on detecting adversarial samples are only valid when the attacker is oblivious to the detection mechanism. In this paper we propose a principled adversarial example detection method that can withstand normconstrained white-box attacks. Inspired by one-versus-the-rest classification, in a K class classification problem, we train K binary classifiers where the i-th binary classifier is used to distinguish between clean data of class i and adversarially perturbed samples of other classes. At test time, we first use a trained classifier to get the predicted label (say k) of the input, and then use the k-th binary classifier to determine whether the input is a clean sample (of class k) or an adversarially perturbed example (of other classes). We further devise a generative approach to detecting/classifying adversarial examples by interpreting each binary classifier as an unnormalized density model of the class-conditional data. We provide comprehensive evaluation of the above adversarial example detection/classification methods, and demonstrate their competitive performances and compelling properties. Code is available at https://github.com/ xuwangyin/GAT-Generative-Adversarial-Training 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd20540a-3965-4489-aa2f-88c78c214c18Cited by top-tier papers9
- Detecting Adversarial Examples Is (Nearly) As Hard As Classifying ThemFlorian TramèrICML 2022 · 82 citations
- DISCO: Adversarial Defense with Local Implicit FunctionsChih-Hui Ho, Nuno VasconcelosNeurIPS 2022 · 65 citations
- Adversarial Attacks on Black Box Video Classifiers: Leveraging the Power of Geometric TransformationsShasha Li, Abhishek Aich, Shitong Zhu, M. Salman Asif et al.NeurIPS 2021 · 50 citations
- Securely Fine-tuning Pre-trained Encoders Against Adversarial ExamplesZiqi Zhou, Minghui Li, Wei Liu, Shengshan Hu et al.S&P 2024 · 23 citations
- Synergy-of-Experts: Collaborate to Improve Adversarial RobustnessSen Cui, Jingfeng Zhang, Jian Liang, Bo Han et al.NeurIPS 2022 · 12 citations
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
Related papers
- AdvFlow: Inconspicuous Black-box Adversarial Attacks using Normalizing FlowsHadi Mohaghegh Dolatabadi, Sarah M. Erfani, Christopher LeckieNeurIPS 2020 · 75 citations
- Adversarial Defense via Learning to Generate Diverse AttacksYunseok Jang, Tianchen Zhao, Seunghoon Hong, Honglak LeeICCV 2019 · 88 citations
- Detecting Adversarial Samples Using Influence Functions and Nearest NeighborsGilad Cohen, Guillermo Sapiro, Raja GiryesCVPR 2020
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- Adversarial Defense by Restricting the Hidden Space of Deep Neural NetworksAamir Mustafa, Salman H. Khan, Munawar Hayat, Roland Goecke et al.ICCV 2019 · 160 citations
