Proper Network Interpretability Helps Adversarial Robustness in Classification
Akhilan Boopathy, Sijia Liu, Gaoyuan Zhang, Cynthia Liu, Pin-Yu Chen, Shiyu Chang, Luca Daniel
摘要
Recent works have empirically shown that there exist adversarial examples that can be hidden from neural network interpretability (namely, making network interpretation maps visually similar), or interpretability is itself susceptible to adversarial attacks. In this paper, we theoretically show that with a proper measurement of interpretation, it is actually difficult to prevent prediction-evasion adversarial attacks from causing interpretation discrepancy, as confirmed by experiments on MNIST, CIFAR-10 and Restricted ImageNet. Spurred by that, we develop an interpretability-aware defensive scheme built only on promoting robust interpretation (without the need for resorting to adversarial loss minimization). We show that our defense achieves both robust classification and robust interpretation, outperforming state-of-the-art adversarial training methods against attacks of large perturbation in particular.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- When does Contrastive Learning Preserve Adversarial Robustness from Pretraining to Finetuning?Lijie Fan, Sijia Liu, Pin-Yu Chen, Gaoyuan Zhang 等NeurIPS 2021 · 被引用 147 次
- Fight Fire with Fire: Towards Robust Recommender Systems via Adversarial Poisoning TrainingChenwang Wu, Defu Lian, Yong Ge, Zhihao Zhu 等SIGIR 2021 · 被引用 47 次
- Adversarial Training for Improving Model Robustness? Look at Both Prediction and InterpretationHanjie Chen, Yangfeng JiAAAI 2022 · 被引用 31 次
- Self-Interpretable Model with Transformation Equivariant InterpretationYipei Wang, Xiaoqian WangNeurIPS 2021 · 被引用 31 次
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
它引用的顶会 Paper3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji 等USENIX Security 2020
相关 Paper
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 被引用 23 次
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
- Fooling Network Interpretation in Image ClassificationAkshayvarun Subramanya, Vipin Pillai, Hamed PirsiavashICCV 2019 · 被引用 68 次
- Nasty Adversarial Training: A Probability Sparsity Perspective for Robustness EnhancementYuhang Zhou, Zhongyun Hua, Zhaoquan Gu, Keke Tang 等ICLR 2026
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
