Fooling Network Interpretation in Image Classification
Akshayvarun Subramanya, Vipin Pillai, Hamed Pirsiavash
摘要
Deep neural networks have been shown to be fooled rather easily using adversarial attack algorithms. Practical methods such as adversarial patches have been shown to be extremely effective in causing misclassification. However, these patches are highlighted using standard network interpretation algorithms, thus revealing the identity of the adversary. We show that it is possible to create adversarial patches which not only fool the prediction, but also change what we interpret regarding the cause of the prediction. Moreover, we introduce our attack as a controlled setting to measure the accuracy of interpretation algorithms. We show this using extensive experiments for Grad-CAM interpretation that transfers to occluding patch interpretation as well. We believe our algorithms can facilitate developing more robust network interpretation tools that truly explain the network's underlying decision making process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Explainable Models with Consistent InterpretationsVipin Pillai, Hamed PirsiavashAAAI 2021 · 被引用 46 次
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 被引用 22 次
- Data Poisoning Attacks Against Outcome Interpretations of Predictive ModelsHengtong Zhang, Jing Gao, Lu SuKDD 2021 · 被引用 21 次
- Saliency-Aware Neural Architecture SearchRamtin Hosseini, Pengtao XieNeurIPS 2022 · 被引用 16 次
它引用的顶会 Paper1
相关 Paper
- Adversarial Attacks on the Interpretation of Neuron Activation MaximizationGéraldin Nanfack, Alexander Fulleringer, Jonathan Marty, Michael Eickenberg 等AAAI 2024 · 被引用 13 次
- Explaining Local, Global, And Higher-Order Interactions In Deep LearningSamuel Lerman, Charles Venuto, Henry A. Kautz, Chenliang XuICCV 2021 · 被引用 13 次
- Robust Feature-Level Adversaries are Interpretability ToolsStephen Casper, Max Nadeau, Dylan Hadfield-Menell, Gabriel KreimanNeurIPS 2022 · 被引用 34 次
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
- Interpretability Based Neural Network RepairZuohui Chen, Jun Zhou, Youcheng Sun, Jingyi Wang 等ISSTA 2024 · 被引用 3 次
