Backdoor Attacks on the DNN Interpretation System
Shihong Fang, Anna Choromanska
Abstract
Interpretability is crucial to understand the inner workings of deep neural networks (DNNs). Many interpretation methods help to understand the decision-making of DNNs by generating saliency maps that highlight parts of the input image that contribute the most to the prediction made by the DNN. In this paper we design a backdoor attack that alters the saliency map produced by the network for an input image with a specific trigger pattern while not losing the prediction performance significantly. The saliency maps are incorporated in the penalty term of the objective function that is used to train a deep model and its influence on model training is conditioned upon the presence of a trigger. We design two types of attacks: a targeted attack that enforces a specific modification of the saliency map and a non-targeted attack when the importance scores of the top pixels from the original saliency map are significantly reduced. We perform empirical evaluations of the proposed backdoor attacks on gradient-based interpretation methods, Grad-CAM and SimpleGrad, and a gradient-free scheme, Vi-sualBackProp, for a variety of deep learning architectures. We show that our attacks constitute a serious security threat to the reliability of the interpretation methods when deploying models developed by untrusted sources. We furthermore show that existing backdoor defense mechanisms are ineffective in detecting our attacks. Finally, we demonstrate that the proposed methodology can be used in an inverted setting, where the correct saliency map can be obtained only in the presence of a trigger (key), effectively making the interpretation system available only to selected users.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a86abc9-4ade-4645-acc5-0f35af16d6a5Cited by top-tier papers9
- Don't trust your eyes: on the (un)reliability of feature visualizationsRobert Geirhos, Roland S. Zimmermann, Blair L. Bilodeau, Wieland Brendel et al.ICML 2024 · 38 citations
- Baffle: Hiding Backdoors in Offline Reinforcement Learning DatasetsChen Gong, Zhou Yang, Yunpeng Bai, Junda He et al.S&P 2024 · 28 citations
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Data Poisoning Attacks Against Outcome Interpretations of Predictive ModelsHengtong Zhang, Jing Gao, Lu SuKDD 2021 · 21 citations
- Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable TriggerYi Yu, Song Xia, Xun Lin, Wenhan Yang et al.AAAI 2025 · 15 citations
Builds on10
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li et al.ICCV 2021 · 639 citations
Related papers
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 62 citations
- Black-box Detection of Backdoor Attacks with Limited Information and DataYinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang et al.ICCV 2021 · 128 citations
- Fooling Network Interpretation in Image ClassificationAkshayvarun Subramanya, Vipin Pillai, Hamed PirsiavashICCV 2019 · 68 citations
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent DiffusionHossein Mirzaei, Zeinab Taghavi, Sepehr Rezaee, Masoud Hadi et al.ICCV 2025 · 3 citations
