Adversarial Defense by Restricting the Hidden Space of Deep Neural Networks
Aamir Mustafa, Salman H. Khan, Munawar Hayat, Roland Goecke, Jianbing Shen, Ling Shao
摘要
Deep neural networks are vulnerable to adversarial attacks, which can fool them by adding minuscule perturbations to the input images. The robustness of existing defenses suffers greatly under white-box attack settings, where an adversary has full knowledge about the network and can iterate several times to find strong perturbations. We observe that the main reason for the existence of such perturbations is the close proximity of different class samples in the learned feature space. This allows model decisions to be totally changed by adding an imperceptible perturbation in the inputs. To counter this, we propose to class-wise disentangle the intermediate feature representations of deep networks. Specifically, we force the features for each class to lie inside a convex polytope that is maximally separated from the polytopes of other classes. In this manner, the network is forced to learn distinct and distant decision regions for each class. We observe that this simple constraint on the features greatly enhances the robustness of learned models, even against the strongest white-box attacks, without degrading the classification performance on clean images. We report extensive evaluations in both black-box and whitebox attack scenarios and show significant gains in comparison to state-of-the art defenses 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Rethinking Softmax Cross-Entropy Loss for Adversarial RobustnessTianyu Pang, Kun Xu, Yinpeng Dong, Chao Du 等ICLR 2020 · 被引用 176 次
- Gaussian Affinity for Max-Margin Class Imbalanced LearningMunawar Hayat, Salman H. Khan, Syed Waqas Zamir, Jianbing Shen 等ICCV 2019 · 被引用 71 次
- DISCO: Adversarial Defense with Local Implicit FunctionsChih-Hui Ho, Nuno VasconcelosNeurIPS 2022 · 被引用 65 次
- Composite Adversarial AttacksXiaofeng Mao, Yuefeng Chen, Shuhui Wang, Hang Su 等AAAI 2021 · 被引用 60 次
它引用的顶会 Paper2
相关 Paper
- LAFEAT: Piercing Through Adversarial Defenses With Latent FeaturesYunrui Yu, Xitong Gao, Cheng-Zhong XuCVPR 2021
- Class-Disentanglement and Applications in Adversarial Detection and DefenseKaiwen Yang, Tianyi Zhou, Yonggang Zhang, Xinmei Tian 等NeurIPS 2021 · 被引用 49 次
- Understanding the Robustness of Randomized Feature Defense Against Query-Based Adversarial AttacksNguyen Hung-Quang, Yingjie Lao, Tung Pham, Kok-Seng Wong 等ICLR 2024 · 被引用 3 次
- Adversarial Feature DesensitizationPouya Bashivan, Reza Bayat, Adam Ibrahim, Kartik Ahuja 等NeurIPS 2021 · 被引用 22 次
- Eliminating Adversarial Noise via Information Discard and Robust Representation RestorationDawei Zhou, Yukun Chen, Nannan Wang, Decheng Liu 等ICML 2023 · 被引用 10 次
