Adversarial Feature Desensitization
Pouya Bashivan, Reza Bayat, Adam Ibrahim, Kartik Ahuja, Mojtaba Faramarzi, Touraj Laleh, Blake A. Richards, Irina Rish
摘要
Neural networks are known to be vulnerable to adversarial attacks -- slight but carefully constructed perturbations of the inputs which can drastically impair the network's performance. Many defense methods have been proposed for improving robustness of deep networks by training them on adversarially perturbed inputs. However, these models often remain vulnerable to new types of attacks not seen during training, and even to slightly stronger versions of previously seen attacks. In this work, we propose a novel approach to adversarial robustness, which builds upon the insights from the domain adaptation field. Our method, called Adversarial Feature Desensitization (AFD), aims at learning features that are invariant towards adversarial perturbations of the inputs. This is achieved through a game where we learn features that are both predictive and robust (insensitive to adversarial attacks), i.e. cannot be used to discriminate between natural and adversarial data. Empirical results on several benchmarks demonstrate the effectiveness of the proposed approach against a wide range of attack types and attack strengths. Our code is available at https://github.com/BashivanLab/afd.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- On the Adversarial Robustness of Out-of-distribution Generalization ModelsXin Zou, Weiwei LiuNeurIPS 2023 · 被引用 10 次
- Towards Better Robustness against Common Corruptions for Unsupervised Domain AdaptationZhiqiang Gao, Kaizhu Huang, Rui Zhang, Dawei Liu 等ICCV 2023 · 被引用 8 次
- Attack-free Evaluating and Enhancing Adversarial Robustness on Categorical DataYujun Zhou, Yufei Han, Haomin Zhuang, Hongyan Bao 等ICML 2024 · 被引用 2 次
它引用的顶会 Paper12
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 被引用 662 次
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal 等ICLR 2020 · 被引用 384 次
- Geometry-aware Instance-reweighted Adversarial TrainingJingfeng Zhang, Jianing Zhu, Gang Niu, Bo Han 等ICLR 2021 · 被引用 316 次
相关 Paper
- Towards Defending against Adversarial Examples via Attack-Invariant FeaturesDawei Zhou, Tongliang Liu, Bo Han, Nannan Wang 等ICML 2021 · 被引用 55 次
- Adversarial Invariant LearningNanyang Ye, Jingxuan Tang, Huayu Deng, Xiao-Yun Zhou 等CVPR 2021
- Attribute-Guided Adversarial Training for Robustness to Natural PerturbationsTejas Gokhale, Rushil Anirudh, Bhavya Kailkhura, Jayaraman J. Thiagarajan 等AAAI 2021 · 被引用 42 次
- Learn2Perturb: An End-to-End Feature Perturbation Learning to Improve Adversarial RobustnessAhmadreza Jeddi, Mohammad Javad Shafiee, Michelle Karg, Christian Scharfenberger 等CVPR 2020
- Adversarial Defense via Learning to Generate Diverse AttacksYunseok Jang, Tianchen Zhao, Seunghoon Hong, Honglak LeeICCV 2019 · 被引用 88 次
