Improving Adversarial Robustness via Mutual Information Estimation
Dawei Zhou, Nannan Wang, Xinbo Gao, Bo Han, Xiaoyu Wang, Yibing Zhan, Tongliang Liu
摘要
Deep neural networks (DNNs) are found to be vulnerable to adversarial noise. They are typically misled by adversarial samples to make wrong predictions. To alleviate this negative effect, in this paper, we investigate the dependence between outputs of the target model and input adversarial samples from the perspective of information theory, and propose an adversarial defense method. Specifically, we first measure the dependence by estimating the mutual information (MI) between outputs and the natural patterns of inputs (called natural MI) and MI between outputs and the adversarial patterns of inputs (called adversarial MI), respectively. We find that adversarial samples usually have larger adversarial MI and smaller natural MI compared with those w.r.t. natural samples. Motivated by this observation, we propose to enhance the adversarial robustness by maximizing the natural MI and minimizing the adversarial MI during the training process. In this way, the target model is expected to pay more attention to the natural pattern that contains objective semantics. Empirical evaluations demonstrate that our method could effectively improve the adversarial accuracy against multiple attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Multimodal Variational Auto-encoder based Audio-Visual SegmentationYuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong 等ICCV 2023 · 被引用 57 次
- Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context LearningZhuo Huang, Chang Liu, Yinpeng Dong, Hang Su 等ICML 2024 · 被引用 31 次
- Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples RegularizationRunqi Lin, Chaojian Yu, Tongliang LiuNeurIPS 2023 · 被引用 25 次
- Phase-aware Adversarial Defense for Improving Adversarial RobustnessDawei Zhou, Nannan Wang, Heng Yang, Xinbo Gao 等ICML 2023 · 被引用 14 次
- Harnessing Out-Of-Distribution Examples via Augmenting Content and StyleZhuo Huang, Xiaobo Xia, Li Shen, Bo Han 等ICLR 2023 · 被引用 10 次
它引用的顶会 Paper14
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
相关 Paper
- Modeling Adversarial Noise for Adversarial TrainingDawei Zhou, Nannan Wang, Bo Han, Tongliang LiuICML 2022 · 被引用 20 次
- Phase and Amplitude-aware Prompting for Enhancing Adversarial RobustnessYibo Xu, Dawei Zhou, Decheng Liu, Nannan WangICML 2025
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 被引用 38 次
- Removing Adversarial Noise in Class Activation Feature SpaceDawei Zhou, Nannan Wang, Chunlei Peng, Xinbo Gao 等ICCV 2021 · 被引用 37 次
- The Enemy of My Enemy is My Friend: Exploring Inverse Adversaries for Improving Adversarial TrainingJunhao Dong, Seyed-Mohsen Moosavi-Dezfooli, Jianhuang Lai, Xiaohua XieCVPR 2023
