What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural Backdoors
Yi-Shan Lin, Wen-Chuan Lee, Z. Berkay Celik
Abstract
EXplainable AI (XAI) methods have been proposed to interpret how a deep neural network predicts inputs through model saliency explanations that highlight the input parts deemed important to arrive at a decision for a specific target. However, it remains challenging to quantify the correctness of their interpretability as current evaluation approaches either require subjective input from humans or incur high computation cost with automated evaluation. In this paper, we propose backdoor trigger patterns--hidden malicious functionalities that cause misclassification--to automate the evaluation of saliency explanations. Our key observation is that triggers provide ground truth for inputs to evaluate whether the regions identified by an XAI method are truly relevant to its output. Since backdoor triggers are the most important features that cause deliberate misclassification, a robust XAI method should reveal their presence at inference time. We introduce three complementary metrics for the systematic evaluation of explanations that an XAI method generates. We evaluate seven state-of-the-art model-free and model-specific post-hoc methods through 36 models trojaned with specifically crafted triggers using color, shape, texture, location, and size. We found six methods that use local explanation and feature relevance fail to completely highlight trigger regions, and only a model-free approach can uncover the entire trigger region. We made our code available at https://github.com/yslin013/evalxai.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71f7edb1-2789-4f82-9fb4-d41f5c2bcd52Cited by top-tier papers8
- Rickrolling the Artist: Injecting Backdoors into Text Encoders for Text-to-Image SynthesisLukas Struppek, Dominik Hintersdorf, Kristian KerstingICCV 2023 · 65 citations
- Towards Reliable and Efficient Backdoor Trigger Inversion via Decoupling Benign FeaturesXiong Xu, Kunzhe Huang, Yiming Li, Zhan Qin et al.ICLR 2024 · 59 citations
- DARE: Disentanglement-Augmented Rationale ExtractionLinan Yue, Qi Liu, Yichao Du, Yanqing An et al.NeurIPS 2022 · 24 citations
- Towards Faithful XAI Evaluation via Generalization-Limited Backdoor WatermarkMengxi Ya, Yiming Li, Tao Dai, Bin Wang et al.ICLR 2024 · 19 citations
- Eye into AI: Evaluating the Interpretability of Explainable AI Techniques through a Game with a PurposeKatelyn Morrison, Mayank Jain, Jessica Hammer, Adam PererCSCW 2023 · 11 citations
Builds on4
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Latent Backdoor Attacks on Deep Neural NetworksYuanshun Yao, Huiying Li, Haitao Zheng, Ben Y. ZhaoCCS 2019 · 465 citations
- An Embarrassingly Simple Approach for Trojan Attack in Deep Neural NetworksRuixiang Tang, Mengnan Du, Ninghao Liu, Fan Yang et al.KDD 2020 · 164 citations
Related papers
- Backdoor Attacks on the DNN Interpretation SystemShihong Fang, Anna ChoromanskaAAAI 2022 · 22 citations
- Xplain: Analyzing Invisible Correlations in Model ExplanationKavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza SadeghiUSENIX Security 2024
- Disguising Attacks with Explanation-Aware BackdoorsMaximilian Noppel, Lukas Peter, Christian WressneggerS&P 2023
- UNICORN: A Unified Backdoor Trigger Inversion FrameworkZhenting Wang, Kai Mei, Juan Zhai, Shiqing MaICLR 2023 · 7 citations
- Deep Feature Space Trojan Attack of Neural Networks by Controlled DetoxificationSiyuan Cheng, Yingqi Liu, Shiqing Ma, Xiangyu ZhangAAAI 2021 · 191 citations
