Removing Adversarial Noise in Class Activation Feature Space
Dawei Zhou, Nannan Wang, Chunlei Peng, Xinbo Gao, Xiaoyu Wang, Jun Yu, Tongliang Liu
Abstract
Deep neural networks (DNNs) are vulnerable to adversarial noise. Pre-processing based defenses could largely remove adversarial noise by processing inputs. However, they are typically affected by the error amplification effect, especially in the front of continuously evolving attacks. To solve this problem, in this paper, we propose to remove adversarial noise by implementing a self-supervised adversarial training mechanism in a class activation feature space. To be specific, we first maximize the disruptions to class activation features of natural examples to craft adversarial examples. Then, we train a denoising model to minimize the distances between the adversarial examples and the natural examples in the class activation feature space. Empirical evaluations demonstrate that our method could significantly enhance adversarial robustness in comparison to previous state-of-the-art approaches, especially against unseen adversarial attacks and adaptive attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Improving Adversarial Robustness via Mutual Information EstimationDawei Zhou, Nannan Wang, Xinbo Gao, Bo Han et al.ICML 2022 · 23 citations
- Modeling Adversarial Noise for Adversarial TrainingDawei Zhou, Nannan Wang, Bo Han, Tongliang LiuICML 2022 · 20 citations
- Robust Overfitting Does Matter: Test-Time Adversarial Purification with FGSMLinyu Tang, Lei ZhangCVPR 2024 · 13 citations
- Improving Accuracy-robustness Trade-off via Pixel Reweighted Adversarial TrainingJiacheng Zhang, Feng Liu, Dawei Zhou, Jingfeng Zhang et al.ICML 2024 · 9 citations
- Rapid Plug-in DefendersKai Wu, Yujian Betterest Li, Jian Lou, Xiaoyu Zhang et al.NeurIPS 2024 · 3 citations
Builds on8
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
Related papers
- Towards Defending against Adversarial Examples via Attack-Invariant FeaturesDawei Zhou, Tongliang Liu, Bo Han, Nannan Wang et al.ICML 2021 · 55 citations
- Pre-trained Adversarial PerturbationsYuanhao Ban, Yinpeng DongNeurIPS 2022 · 37 citations
- A Self-supervised Approach for Adversarial RobustnessMuzammal Naseer, Salman H. Khan, Munawar Hayat, Fahad Shahbaz Khan et al.CVPR 2020
- Improving Adversarial Robustness via Channel-wise Activation SuppressingYang Bai, Yuyuan Zeng, Yong Jiang, Shu-Tao Xia et al.ICLR 2021 · 59 citations
- Online Adversarial Purification based on Self-supervised LearningChanghao Shi, Chester Holtz, Gal MishneICLR 2021 · 63 citations
