On the Robustness of Removal-Based Feature Attributions
Chris Lin, Ian Covert, Su-In Lee
Abstract
To explain predictions made by complex machine learning models, many feature attribution methods have been developed that assign importance scores to input features. Some recent work challenges the robustness of these methods by showing that they are sensitive to input and model perturbations, while other work addresses this issue by proposing robust attribution methods. However, previous work on attribution robustness has focused primarily on gradient-based feature attributions, whereas the robustness of removal-based attribution methods is not currently well understood. To bridge this gap, we theoretically characterize the robustness properties of removal-based feature attributions. Specifically, we provide a unified analysis of such methods and derive upper bounds for the difference between intact and perturbed attributions, under settings of both input and model perturbations. Our empirical results on synthetic and real-world data validate our theoretical results and demonstrate their practical implications, including the ability to increase attribution robustness by improving the model's Lipschitz regularity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5737fd1b-0b99-4024-9b0e-d9bb1adbc782Cited by top-tier papers9
- Stochastic Amortization: A Unified Approach to Accelerate Feature and Data AttributionIan Covert, Chanwoo Kim, Su-In Lee, James Y. Zou et al.NeurIPS 2024 · 25 citations
- Probabilistic Stability Guarantees for Feature AttributionsHelen Jin, Anton Xue, Weiqiu You, Surbhi Goel et al.NeurIPS 2025 · 12 citations
- Provably Better Explanations with Optimized Aggregation of Feature AttributionsThomas Decker, Ananta R. Bhattarai, Jindong Gu, Volker Tresp et al.ICML 2024 · 7 citations
- Improving Perturbation-based Explanations by Understanding the Role of Uncertainty CalibrationThomas Decker, Volker Tresp, Florian BuettnerNeurIPS 2025 · 3 citations
- Missingness Bias Calibration in Feature Attribution ExplanationsShailesh Sridhar, Anton Xue, Eric WongICLR 2026
Builds on10
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 799 citations
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 476 citations
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya FeigeNeurIPS 2020 · 246 citations
- Shapley explainability on the data manifoldChristopher Frye, Damien de Mijolla, Tom Begley, Laurence Cowton et al.ICLR 2021 · 125 citations
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller et al.ICML 2020 · 104 citations
Related papers
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- Stability Guarantees for Feature Attributions with Multiplicative SmoothingAnton Xue, Rajeev Alur, Eric WongNeurIPS 2023 · 18 citations
- Rethinking Robustness of Model AttributionsSandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N. BalasubramanianAAAI 2024 · 2 citations
- Evaluations and Methods for Explanation through Robustness AnalysisCheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Kumar Ravikumar et al.ICLR 2021 · 68 citations
- Re-calibrating Feature Attributions for Model InterpretationPeiyu Yang, Naveed Akhtar, Zeyi Wen, Mubarak Shah et al.ICLR 2023
