"Is your explanation stable?": A Robustness Evaluation Framework for Feature Attribution
Yuyou Gan, Yuhao Mao, Xuhong Zhang, Shouling Ji, Yuwen Pu, Meng Han, Jianwei Yin, Ting Wang
Abstract
Neural networks have become increasingly popular. Nevertheless, understanding their decision process turns out to be complicated. One vital method to explain a models' behavior is feature attribution, i.e., attributing its decision to pivotal features. Although many algorithms are proposed, most of them aim to improve the faithfulness (fidelity) to the model. However, the real environment contains many random noises, which may cause the feature attribution maps to be greatly perturbed for similar images. More seriously, recent works show that explanation algorithms are vulnerable to adversarial attacks, generating the same explanation for a maliciously perturbed input. All of these make the explanation hard to trust in real scenarios, especially in security-critical applications. To bridge this gap, we propose Median Test for Feature Attribution (MeTFA) to quantify the uncertainty and increase the stability of explanation algorithms with theoretical guarantees. MeTFA is method-agnostic, i.e., it can be applied to any feature attribution method. MeTFA has the following two functions: (1) examine whether one feature is significantly important or unimportant and generate a MeTFA-significant map to visualize the results; (2) compute the confidence interval of a feature attribution score and generate a MeTFA-smoothed map to increase the stability of the explanation. Extensive experiments show that MeTFA improves the visual quality of explanations and significantly reduces the instability while maintaining the faithfulness of the original method. To quantitatively evaluate MeTFA's faithfulness and stability, we further propose several robust faithfulness metrics, which can evaluate * Yuyou Gan and Yuhao Mao contributed equally. Xuhong Zhang and Shouling Ji are the corresponding authors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Probabilistic Stability Guarantees for Feature AttributionsHelen Jin, Anton Xue, Weiqiu You, Surbhi Goel et al.NeurIPS 2025 · 12 citations
- Good-looking but Lacking Faithfulness: Understanding Local Explanation Methods through Trend-based TestingJinwen He, Kai Chen, Guozhu Meng, Jiangshan Zhang et al.CCS 2023 · 2 citations
- Everybody's Got ML, Tell Me What Else You Have: Practitioners' Perception of ML-Based Security Tools and ExplanationsJaron Mink, Hadjer Benkraouda, Limin Yang, Arridhana Ciptadi et al.S&P 2023
Builds on15
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su et al.CCS 2018 · 336 citations
- Visualizing Deep Networks by Optimizing with Integrated GradientsZhongang Qi, Saeed Khorram, Fuxin LiAAAI 2020 · 149 citations
- Towards the Unification and Robustness of Perturbation and Gradient Based ExplanationsSushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay et al.ICML 2021 · 71 citations
- NeuronFair: Interpretable White-Box Fairness Testing through Biased Neuron IdentificationHaibin Zheng, Zhiqing Chen, Tianyu Du, Xuhong Zhang et al.ICSE 2022 · 58 citations
- A Novel Visual Interpretability for Deep Neural Networks by Optimizing Activation Maps with PerturbationQing-Long Zhang, Lu Rao, Yubin YangAAAI 2021 · 26 citations
Related papers
- Stability Guarantees for Feature Attributions with Multiplicative SmoothingAnton Xue, Rajeev Alur, Eric WongNeurIPS 2023 · 18 citations
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- Provably Better Explanations with Optimized Aggregation of Feature AttributionsThomas Decker, Ananta R. Bhattarai, Jindong Gu, Volker Tresp et al.ICML 2024 · 7 citations
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 32 citations
- Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions ExplainabilityJoakim Edin, Andreas Geert Motzfeldt, Casper L. Christensen, Tuukka Ruotsalo et al.ACL 2025
