Improving Adversarial Robustness via Mutual Information Estimation
Dawei Zhou, Nannan Wang, Xinbo Gao, Bo Han, Xiaoyu Wang, Yibing Zhan, Tongliang Liu
Abstract
Deep neural networks (DNNs) are found to be vulnerable to adversarial noise. They are typically misled by adversarial samples to make wrong predictions. To alleviate this negative effect, in this paper, we investigate the dependence between outputs of the target model and input adversarial samples from the perspective of information theory, and propose an adversarial defense method. Specifically, we first measure the dependence by estimating the mutual information (MI) between outputs and the natural patterns of inputs (called natural MI) and MI between outputs and the adversarial patterns of inputs (called adversarial MI), respectively. We find that adversarial samples usually have larger adversarial MI and smaller natural MI compared with those w.r.t. natural samples. Motivated by this observation, we propose to enhance the adversarial robustness by maximizing the natural MI and minimizing the adversarial MI during the training process. In this way, the target model is expected to pay more attention to the natural pattern that contains objective semantics. Empirical evaluations demonstrate that our method could effectively improve the adversarial accuracy against multiple attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a21f5565-bfee-473a-b046-662deebd7befCited by top-tier papers11
- Multimodal Variational Auto-encoder based Audio-Visual SegmentationYuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong et al.ICCV 2023 · 57 citations
- Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context LearningZhuo Huang, Chang Liu, Yinpeng Dong, Hang Su et al.ICML 2024 · 31 citations
- Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples RegularizationRunqi Lin, Chaojian Yu, Tongliang LiuNeurIPS 2023 · 25 citations
- Phase-aware Adversarial Defense for Improving Adversarial RobustnessDawei Zhou, Nannan Wang, Heng Yang, Xinbo Gao et al.ICML 2023 · 14 citations
- Harnessing Out-Of-Distribution Examples via Augmenting Content and StyleZhuo Huang, Xiaobo Xia, Li Shen, Bo Han et al.ICLR 2023 · 10 citations
Builds on14
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
Related papers
- Modeling Adversarial Noise for Adversarial TrainingDawei Zhou, Nannan Wang, Bo Han, Tongliang LiuICML 2022 · 20 citations
- Phase and Amplitude-aware Prompting for Enhancing Adversarial RobustnessYibo Xu, Dawei Zhou, Decheng Liu, Nannan WangICML 2025
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 38 citations
- Removing Adversarial Noise in Class Activation Feature SpaceDawei Zhou, Nannan Wang, Chunlei Peng, Xinbo Gao et al.ICCV 2021 · 37 citations
- The Enemy of My Enemy is My Friend: Exploring Inverse Adversaries for Improving Adversarial TrainingJunhao Dong, Seyed-Mohsen Moosavi-Dezfooli, Jianhuang Lai, Xiaohua XieCVPR 2023
