Towards Defending against Adversarial Examples via Attack-Invariant Features
Dawei Zhou, Tongliang Liu, Bo Han, Nannan Wang, Chunlei Peng, Xinbo Gao
Abstract
Deep neural networks (DNNs) are vulnerable to adversarial noise. Their adversarial robustness can be improved by exploiting adversarial examples. However, given the continuously evolving attacks, models trained on seen types of adversarial examples generally cannot generalize well to unseen types of adversarial examples. To solve this problem, in this paper, we propose to remove adversarial noise by learning generalizable invariant features across attacks which maintain semantic classification information. Specifically, we introduce an adversarial feature learning mechanism to disentangle invariant features from adversarial noise. A normalization term has been proposed in the encoded space of the attack-invariant features to address the bias issue between the seen and unseen types of attacks. Empirical evaluations demonstrate that our method could provide better protection in comparison to previous state-of-the-art approaches, especially against unseen types of attacks and adaptive attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1dd8e3c4-9ae5-4b1b-a7e3-1ff6b294efccCited by top-tier papers14
- Probabilistic Margins for Instance Reweighting in Adversarial TrainingQizhou Wang, Feng Liu, Bo Han, Tongliang Liu et al.NeurIPS 2021 · 84 citations
- Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object DetectionLongrong Yang, Xianpan Zhou, Xuewei Li, Liang Qiao et al.ICCV 2023 · 51 citations
- Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context LearningZhuo Huang, Chang Liu, Yinpeng Dong, Hang Su et al.ICML 2024 · 31 citations
- Improving Adversarial Robustness via Mutual Information EstimationDawei Zhou, Nannan Wang, Xinbo Gao, Bo Han et al.ICML 2022 · 23 citations
- Modeling Adversarial Noise for Adversarial TrainingDawei Zhou, Nannan Wang, Bo Han, Tongliang LiuICML 2022 · 20 citations
Builds on8
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Part-dependent Label Noise: Towards Instance-dependent Label NoiseXiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang et al.NeurIPS 2020 · 329 citations
Related papers
- Removing Adversarial Noise in Class Activation Feature SpaceDawei Zhou, Nannan Wang, Chunlei Peng, Xinbo Gao et al.ICCV 2021 · 37 citations
- Adversarial Feature DesensitizationPouya Bashivan, Reza Bayat, Adam Ibrahim, Kartik Ahuja et al.NeurIPS 2021 · 22 citations
- Adversarial Invariant LearningNanyang Ye, Jingxuan Tang, Huayu Deng, Xiao-Yun Zhou et al.CVPR 2021
- StyLess: Boosting the Transferability of Adversarial ExamplesKaisheng Liang, Bin XiaoCVPR 2023
- Domain Generalization by Learning and Removing Domain-specific FeaturesYu Ding, Lei Wang, Bin Liang, Shuming Liang et al.NeurIPS 2022 · 75 citations
