Robust Explanation for Free or At the Cost of Faithfulness
Zeren Tan, Yang Tian
Abstract
Devoted to interpreting the explicit behaviors of machine learning models, explanation methods can identify implicit characteristics of models to improve trustworthiness. However, explanation methods are shown as vulnerable to adversarial perturbations, implying security concerns in highstakes domains. In this paper, we investigate when robust explanations are necessary and what they cost. We prove that the robustness of explanations is determined by the robustness of the model to be explained. Therefore, we can have robust explanations for free for a robust model. To have robust explanations for a non-robust model, composing the original model with a kernel is proved as an effective way that returns strictly more robust explanations. Nevertheless, we argue that this also incurs a robustness-faithfulness trade-off, i.e., contrary to common expectations, an explanation method may also become less faithful when it becomes more robust. This argument holds for any model. We are the first to introduce this trade-off and theoretically prove its existence for SmoothGrad. Theoretical findings are verified by empirical evidence on six state-of-the-art explanation methods and four backbones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b03fd085-90ef-4fa6-94d3-e2970d7f7f98Cited by top-tier papers4
- Robust Explanations of Graph Neural Networks via Graph CurvaturesYazheng Liu, Xi Zhang, Sihong Xie, Hui XiongNeurIPS 2025 · 1 citation
- Factor Graph-based Interpretable Neural NetworksYicong Li, Kuanjiu Zhou, Shuo Yu, Qiang Zhang et al.ICLR 2025
- Pixel-level Certified Explanations via Randomized SmoothingAlaa Anani, Tobias Lorenz, Mario Fritz, Bernt SchieleICML 2025
- Provably Robust Explainable Graph Neural Networks against Graph Perturbation AttacksJiate Li, Meng Pang, Yun Dong, Jinyuan Jia et al.ICLR 2025
Builds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Counterfactual Explanations Can Be ManipulatedDylan Slack, Anna Hilgard, Himabindu Lakkaraju, Sameer SinghNeurIPS 2021 · 182 citations
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- Towards Robust and Reliable Algorithmic RecourseSohini Upadhyay, Shalmali Joshi, Himabindu LakkarajuNeurIPS 2021 · 145 citations
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller et al.ICML 2020 · 104 citations
Related papers
- Evaluations and Methods for Explanation through Robustness AnalysisCheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Kumar Ravikumar et al.ICLR 2021 · 68 citations
- Use perturbations when learning from explanationsJuyeon Heo, Vihari Piratla, Matthew Wicker, Adrian WellerNeurIPS 2023 · 2 citations
- Towards Automating Model Explanations with Certified Robustness GuaranteesMengdi Huai, Jinduo Liu, Chenglin Miao, Liuyi Yao et al.AAAI 2022 · 16 citations
- Rethinking Robustness of Model AttributionsSandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N. BalasubramanianAAAI 2024 · 2 citations
- Robust Explanation Constraints for Neural NetworksMatthew Wicker, Juyeon Heo, Luca Costabello, Adrian WellerICLR 2023 · 3 citations
