Towards a Unified Game-Theoretic View of Adversarial Perturbations and Robustness
Jie Ren, Die Zhang, Yisen Wang, Lu Chen, Zhanpeng Zhou, Yiting Chen, Xu Cheng, Xin Wang, Meng Zhou, Jie Shi, Quanshi Zhang
摘要
This paper provides a unified view to explain different adversarial attacks and defense methods, i.e. the view of multi-order interactions between input variables of DNNs. Based on the multi-order interaction, we discover that adversarial attacks mainly affect high-order interactions to fool the DNN. Furthermore, we find that the robustness of adversarially trained DNNs comes from category-specific low-order interactions. Our findings provide a potential method to unify adversarial perturbations and robustness, which can explain the existing robustness-boosting methods in a principle way. Besides, our findings also make a revision of previous inaccurate understanding of the shape bias of adversarially learned features. Our code is available online at https://github.com/Jie-Ren/A-Unified-Game-Theoretic-Interpretation-of-Adversarial-Robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Does a Neural Network Really Encode Symbolic Concepts?Mingjie Li, Quanshi ZhangICML 2023 · 被引用 35 次
- Where We Have Arrived in Proving the Emergence of Sparse Interaction Primitives in DNNsQihan Ren, Jiayang Gao, Wen Shen, Quanshi ZhangICLR 2024 · 被引用 23 次
- Towards the Dynamics of a DNN Learning Symbolic InteractionsQihan Ren, Junpeng Zhang, Yang Xu, Yue Xin 等NeurIPS 2024 · 被引用 21 次
- A Simple Yet Effective Strategy to Robustify the Meta Learning ParadigmQi Wang, Yiqin Lv, Yang-He Feng, Zheng Xie 等NeurIPS 2023 · 被引用 17 次
- Feature Attribution with Necessity and Sufficiency via Dual-stage Perturbation Test for Causal ExplanationXuexin Chen, Ruichu Cai, Zhengting Huang, Yuxuan Zhu 等ICML 2024 · 被引用 5 次
它引用的顶会 Paper18
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor 等NeurIPS 2020 · 被引用 506 次
相关 Paper
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu 等ICLR 2021 · 被引用 113 次
- Interpreting Attributions and Interactions of Adversarial AttacksXin Wang, Shuyun Lin, Hao Zhang, Yufei Zhu 等ICCV 2021 · 被引用 20 次
- Interpreting Robustness Proofs of Deep Neural NetworksDebangshu Banerjee, Avaljot Singh, Gagandeep SinghICLR 2024 · 被引用 6 次
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
- Improving Adversarial Robustness via Mutual Information EstimationDawei Zhou, Nannan Wang, Xinbo Gao, Bo Han 等ICML 2022 · 被引用 23 次
