Fooling Network Interpretation in Image Classification
Akshayvarun Subramanya, Vipin Pillai, Hamed Pirsiavash
Abstract
Deep neural networks have been shown to be fooled rather easily using adversarial attack algorithms. Practical methods such as adversarial patches have been shown to be extremely effective in causing misclassification. However, these patches are highlighted using standard network interpretation algorithms, thus revealing the identity of the adversary. We show that it is possible to create adversarial patches which not only fool the prediction, but also change what we interpret regarding the cause of the prediction. Moreover, we introduce our attack as a controlled setting to measure the accuracy of interpretation algorithms. We show this using extensive experiments for Grad-CAM interpretation that transfers to occluding patch interpretation as well. We believe our algorithms can facilitate developing more robust network interpretation tools that truly explain the network's underlying decision making process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a976313-86d6-4bb7-8f05-e747fcddeeb2Cited by top-tier papers13
- Explainable Models with Consistent InterpretationsVipin Pillai, Hamed PirsiavashAAAI 2021 · 46 citations
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 22 citations
- Data Poisoning Attacks Against Outcome Interpretations of Predictive ModelsHengtong Zhang, Jing Gao, Lu SuKDD 2021 · 21 citations
- Saliency-Aware Neural Architecture SearchRamtin Hosseini, Pengtao XieNeurIPS 2022 · 16 citations
Builds on1
Related papers
- Adversarial Attacks on the Interpretation of Neuron Activation MaximizationGéraldin Nanfack, Alexander Fulleringer, Jonathan Marty, Michael Eickenberg et al.AAAI 2024 · 13 citations
- Explaining Local, Global, And Higher-Order Interactions In Deep LearningSamuel Lerman, Charles Venuto, Henry A. Kautz, Chenliang XuICCV 2021 · 13 citations
- Robust Feature-Level Adversaries are Interpretability ToolsStephen Casper, Max Nadeau, Dylan Hadfield-Menell, Gabriel KreimanNeurIPS 2022 · 34 citations
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
- Interpretability Based Neural Network RepairZuohui Chen, Jun Zhou, Youcheng Sun, Jingyi Wang et al.ISSTA 2024 · 3 citations
