Interpreting Attributions and Interactions of Adversarial Attacks
Xin Wang, Shuyun Lin, Hao Zhang, Yufei Zhu, Quanshi Zhang
Abstract
This paper aims to explain adversarial attacks in terms of how adversarial perturbations contribute to the attacking task. We estimate attributions of different image regions to the decrease of the attacking cost based on the Shapley value. We define and quantify interactions among adversarial perturbation pixels, and decompose the entire perturbation map into relatively independent perturbation components. The decomposition of the perturbation map shows that adversarially-trained DNNs have more perturbation components in the foreground than normally-trained DNNs. Moreover, compared to the normally-trained DNN, the adversarially-trained DNN have more components which mainly decrease the score of the true category. Above analyses provide new insights into the understanding of adversarial attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77e38ac7-ee89-4e09-9d3c-8ae7b72d3b4eCited by top-tier papers12
- Discovering and Explaining the Representation Bottleneck of DNNSHuiqi Deng, Qihan Ren, Hao Zhang, Quanshi ZhangICLR 2022 · 73 citations
- Does a Neural Network Really Encode Symbolic Concepts?Mingjie Li, Quanshi ZhangICML 2023 · 35 citations
- Towards a Unified Game-Theoretic View of Adversarial Perturbations and RobustnessJie Ren, Die Zhang, Yisen Wang, Lu Chen et al.NeurIPS 2021 · 27 citations
- Towards Understanding the Generalization of Deepfake Detectors from a Game-Theoretical ViewKelu Yao, Jin Wang, Boyu Diao, Chao LiICCV 2023 · 26 citations
- Visualizing the Emergence of Intermediate Visual Patterns in DNNsMingjie Li, Shaobo Wang, Quanshi ZhangNeurIPS 2021 · 12 citations
Builds on6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu et al.ICLR 2021 · 113 citations
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue et al.ICLR 2020 · 55 citations
- Interpreting and Boosting Dropout from a Game-Theoretic ViewHao Zhang, Sen Li, Yinchao Ma, Mingjie Li et al.ICLR 2021 · 53 citations
Related papers
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- Interpreting Multivariate Shapley Interactions in DNNsHao Zhang, Yichen Xie, Longjie Zheng, Die Zhang et al.AAAI 2021 · 70 citations
- Rethinking and Improving Robustness of Convolutional Neural Networks: a Shapley Value-based Approach in Frequency DomainYiting Chen, Qibing Ren, Junchi YanNeurIPS 2022 · 36 citations
- Building Interpretable Interaction Trees for Deep NLP ModelsDie Zhang, Hao Zhang, Huilin Zhou, Xiaoyi Bao et al.AAAI 2021 · 43 citations
- Enhancing Interpretability for Vision Models via Shapley Value OptimizationKanglong Fan, Yunqiao Yang, Chen MaAAAI 2026
