Robust Explanation for Free or At the Cost of Faithfulness
Zeren Tan, Yang Tian
摘要
Devoted to interpreting the explicit behaviors of machine learning models, explanation methods can identify implicit characteristics of models to improve trustworthiness. However, explanation methods are shown as vulnerable to adversarial perturbations, implying security concerns in highstakes domains. In this paper, we investigate when robust explanations are necessary and what they cost. We prove that the robustness of explanations is determined by the robustness of the model to be explained. Therefore, we can have robust explanations for free for a robust model. To have robust explanations for a non-robust model, composing the original model with a kernel is proved as an effective way that returns strictly more robust explanations. Nevertheless, we argue that this also incurs a robustness-faithfulness trade-off, i.e., contrary to common expectations, an explanation method may also become less faithful when it becomes more robust. This argument holds for any model. We are the first to introduce this trade-off and theoretically prove its existence for SmoothGrad. Theoretical findings are verified by empirical evidence on six state-of-the-art explanation methods and four backbones.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Robust Explanations of Graph Neural Networks via Graph CurvaturesYazheng Liu, Xi Zhang, Sihong Xie, Hui XiongNeurIPS 2025 · 被引用 1 次
- Factor Graph-based Interpretable Neural NetworksYicong Li, Kuanjiu Zhou, Shuo Yu, Qiang Zhang 等ICLR 2025
- Pixel-level Certified Explanations via Randomized SmoothingAlaa Anani, Tobias Lorenz, Mario Fritz, Bernt SchieleICML 2025
- Provably Robust Explainable Graph Neural Networks against Graph Perturbation AttacksJiate Li, Meng Pang, Yun Dong, Jinyuan Jia 等ICLR 2025
它引用的顶会 Paper12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Counterfactual Explanations Can Be ManipulatedDylan Slack, Anna Hilgard, Himabindu Lakkaraju, Sameer SinghNeurIPS 2021 · 被引用 182 次
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu 等ICML 2020 · 被引用 148 次
- Towards Robust and Reliable Algorithmic RecourseSohini Upadhyay, Shalmali Joshi, Himabindu LakkarajuNeurIPS 2021 · 被引用 145 次
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller 等ICML 2020 · 被引用 104 次
相关 Paper
- Evaluations and Methods for Explanation through Robustness AnalysisCheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Kumar Ravikumar 等ICLR 2021 · 被引用 68 次
- Use perturbations when learning from explanationsJuyeon Heo, Vihari Piratla, Matthew Wicker, Adrian WellerNeurIPS 2023 · 被引用 2 次
- Towards Automating Model Explanations with Certified Robustness GuaranteesMengdi Huai, Jinduo Liu, Chenglin Miao, Liuyi Yao 等AAAI 2022 · 被引用 16 次
- Rethinking Robustness of Model AttributionsSandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N. BalasubramanianAAAI 2024 · 被引用 2 次
- Robust Explanation Constraints for Neural NetworksMatthew Wicker, Juyeon Heo, Luca Costabello, Adrian WellerICLR 2023 · 被引用 3 次
