Interpreting Attributions and Interactions of Adversarial Attacks
Xin Wang, Shuyun Lin, Hao Zhang, Yufei Zhu, Quanshi Zhang
摘要
This paper aims to explain adversarial attacks in terms of how adversarial perturbations contribute to the attacking task. We estimate attributions of different image regions to the decrease of the attacking cost based on the Shapley value. We define and quantify interactions among adversarial perturbation pixels, and decompose the entire perturbation map into relatively independent perturbation components. The decomposition of the perturbation map shows that adversarially-trained DNNs have more perturbation components in the foreground than normally-trained DNNs. Moreover, compared to the normally-trained DNN, the adversarially-trained DNN have more components which mainly decrease the score of the true category. Above analyses provide new insights into the understanding of adversarial attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Discovering and Explaining the Representation Bottleneck of DNNSHuiqi Deng, Qihan Ren, Hao Zhang, Quanshi ZhangICLR 2022 · 被引用 73 次
- Does a Neural Network Really Encode Symbolic Concepts?Mingjie Li, Quanshi ZhangICML 2023 · 被引用 35 次
- Towards a Unified Game-Theoretic View of Adversarial Perturbations and RobustnessJie Ren, Die Zhang, Yisen Wang, Lu Chen 等NeurIPS 2021 · 被引用 27 次
- Towards Understanding the Generalization of Deepfake Detectors from a Game-Theoretical ViewKelu Yao, Jin Wang, Boyu Diao, Chao LiICCV 2023 · 被引用 26 次
- Visualizing the Emergence of Intermediate Visual Patterns in DNNsMingjie Li, Shaobo Wang, Quanshi ZhangNeurIPS 2021 · 被引用 12 次
它引用的顶会 Paper6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu 等ICLR 2021 · 被引用 113 次
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue 等ICLR 2020 · 被引用 55 次
- Interpreting and Boosting Dropout from a Game-Theoretic ViewHao Zhang, Sen Li, Yinchao Ma, Mingjie Li 等ICLR 2021 · 被引用 53 次
相关 Paper
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu 等ICML 2020 · 被引用 148 次
- Interpreting Multivariate Shapley Interactions in DNNsHao Zhang, Yichen Xie, Longjie Zheng, Die Zhang 等AAAI 2021 · 被引用 70 次
- Rethinking and Improving Robustness of Convolutional Neural Networks: a Shapley Value-based Approach in Frequency DomainYiting Chen, Qibing Ren, Junchi YanNeurIPS 2022 · 被引用 36 次
- Building Interpretable Interaction Trees for Deep NLP ModelsDie Zhang, Hao Zhang, Huilin Zhou, Xiaoyi Bao 等AAAI 2021 · 被引用 43 次
- Enhancing Interpretability for Vision Models via Shapley Value OptimizationKanglong Fan, Yunqiao Yang, Chen MaAAAI 2026
