Rethinking Robustness of Model Attributions
Sandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N. Balasubramanian
Abstract
For machine learning models to be reliable and trustworthy, their decisions must be interpretable. As these models find increasing use in safety-critical applications, it is important that not just the model predictions but also their explanations (as feature attributions) be robust to small human-imperceptible input perturbations. Recent works have shown that many attribution methods are fragile and have proposed improvements in either these methods or the model training. We observe two main causes for fragile attributions: first, the existing metrics of robustness (e.g., top-k intersection) overpenalize even reasonable local shifts in attribution, thereby making random perturbations to appear as a strong attack, and second, the attribution can be concentrated in a small region even when there are multiple important parts in an image. To rectify this, we propose simple ways to strengthen existing metrics and attribution methods that incorporate locality of pixels in robustness metrics and diversity of pixel locations in attributions. Towards the role of model training in attributional robustness, we empirically observe that adversarially trained models have more robust attributions on smaller datasets, however, this advantage disappears in larger datasets. Code is made available at https://github.com/ksandeshk/LENS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Probabilistic Stability Guarantees for Feature AttributionsHelen Jin, Anton Xue, Weiqiu You, Surbhi Goel et al.NeurIPS 2025 · 12 citations
- Factor Graph-based Interpretable Neural NetworksYicong Li, Kuanjiu Zhou, Shuo Yu, Qiang Zhang et al.ICLR 2025
- Improving Adversarial Robustness of Attribution via Implicit RegularizationAmir Mehrpanah, Matteo Gamba, Hossein AzizpourICML 2026
Builds on6
- Reliable Post hoc Explanations: Modeling Uncertainty in ExplainabilityDylan Slack, Anna Hilgard, Sameer Singh, Himabindu LakkarajuNeurIPS 2021 · 240 citations
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram et al.AAAI 2020 · 204 citations
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 93 citations
- Enhanced Regularizers for Attributional RobustnessAnindya Sarkar, Anirban Sarkar, Vineeth N. BalasubramanianAAAI 2021 · 18 citations
Related papers
- SAM: The Sensitivity of Attribution Methods to HyperparametersNaman Bansal, Chirag Agarwal, Anh NguyenCVPR 2020
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- Pixel-level Certified Explanations via Randomized SmoothingAlaa Anani, Tobias Lorenz, Mario Fritz, Bernt SchieleICML 2025
- On the Robustness of Removal-Based Feature AttributionsChris Lin, Ian Covert, Su-In LeeNeurIPS 2023 · 25 citations
- Towards More Robust Interpretation via Local Gradient AlignmentSunghwan Joo, Seokhyeon Jeong, Juyeon Heo, Adrian Weller et al.AAAI 2023 · 8 citations
