Adversarial Defense by Restricting the Hidden Space of Deep Neural Networks
Aamir Mustafa, Salman H. Khan, Munawar Hayat, Roland Goecke, Jianbing Shen, Ling Shao
Abstract
Deep neural networks are vulnerable to adversarial attacks, which can fool them by adding minuscule perturbations to the input images. The robustness of existing defenses suffers greatly under white-box attack settings, where an adversary has full knowledge about the network and can iterate several times to find strong perturbations. We observe that the main reason for the existence of such perturbations is the close proximity of different class samples in the learned feature space. This allows model decisions to be totally changed by adding an imperceptible perturbation in the inputs. To counter this, we propose to class-wise disentangle the intermediate feature representations of deep networks. Specifically, we force the features for each class to lie inside a convex polytope that is maximally separated from the polytopes of other classes. In this manner, the network is forced to learn distinct and distant decision regions for each class. We observe that this simple constraint on the features greatly enhances the robustness of learned models, even against the strongest white-box attacks, without degrading the classification performance on clean images. We report extensive evaluations in both black-box and whitebox attack scenarios and show significant gains in comparison to state-of-the art defenses 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0426f4e1-5300-4b69-a0f3-27cbc7329acaCited by top-tier papers24
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Rethinking Softmax Cross-Entropy Loss for Adversarial RobustnessTianyu Pang, Kun Xu, Yinpeng Dong, Chao Du et al.ICLR 2020 · 176 citations
- Gaussian Affinity for Max-Margin Class Imbalanced LearningMunawar Hayat, Salman H. Khan, Syed Waqas Zamir, Jianbing Shen et al.ICCV 2019 · 71 citations
- DISCO: Adversarial Defense with Local Implicit FunctionsChih-Hui Ho, Nuno VasconcelosNeurIPS 2022 · 65 citations
- Composite Adversarial AttacksXiaofeng Mao, Yuefeng Chen, Shuhui Wang, Hang Su et al.AAAI 2021 · 60 citations
Builds on2
Related papers
- LAFEAT: Piercing Through Adversarial Defenses With Latent FeaturesYunrui Yu, Xitong Gao, Cheng-Zhong XuCVPR 2021
- Class-Disentanglement and Applications in Adversarial Detection and DefenseKaiwen Yang, Tianyi Zhou, Yonggang Zhang, Xinmei Tian et al.NeurIPS 2021 · 49 citations
- Understanding the Robustness of Randomized Feature Defense Against Query-Based Adversarial AttacksNguyen Hung-Quang, Yingjie Lao, Tung Pham, Kok-Seng Wong et al.ICLR 2024 · 3 citations
- Adversarial Feature DesensitizationPouya Bashivan, Reza Bayat, Adam Ibrahim, Kartik Ahuja et al.NeurIPS 2021 · 22 citations
- Eliminating Adversarial Noise via Information Discard and Robust Representation RestorationDawei Zhou, Yukun Chen, Nannan Wang, Decheng Liu et al.ICML 2023 · 10 citations
