Rethinking Robustness of Model Attributions
Sandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N. Balasubramanian
摘要
For machine learning models to be reliable and trustworthy, their decisions must be interpretable. As these models find increasing use in safety-critical applications, it is important that not just the model predictions but also their explanations (as feature attributions) be robust to small human-imperceptible input perturbations. Recent works have shown that many attribution methods are fragile and have proposed improvements in either these methods or the model training. We observe two main causes for fragile attributions: first, the existing metrics of robustness (e.g., top-k intersection) overpenalize even reasonable local shifts in attribution, thereby making random perturbations to appear as a strong attack, and second, the attribution can be concentrated in a small region even when there are multiple important parts in an image. To rectify this, we propose simple ways to strengthen existing metrics and attribution methods that incorporate locality of pixels in robustness metrics and diversity of pixel locations in attributions. Towards the role of model training in attributional robustness, we empirically observe that adversarially trained models have more robust attributions on smaller datasets, however, this advantage disappears in larger datasets. Code is made available at https://github.com/ksandeshk/LENS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Probabilistic Stability Guarantees for Feature AttributionsHelen Jin, Anton Xue, Weiqiu You, Surbhi Goel 等NeurIPS 2025 · 被引用 12 次
- Factor Graph-based Interpretable Neural NetworksYicong Li, Kuanjiu Zhou, Shuo Yu, Qiang Zhang 等ICLR 2025
- Improving Adversarial Robustness of Attribution via Implicit RegularizationAmir Mehrpanah, Matteo Gamba, Hossein AzizpourICML 2026
它引用的顶会 Paper6
- Reliable Post hoc Explanations: Modeling Uncertainty in ExplainabilityDylan Slack, Anna Hilgard, Sameer Singh, Himabindu LakkarajuNeurIPS 2021 · 被引用 240 次
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram 等AAAI 2020 · 被引用 204 次
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu 等ICML 2020 · 被引用 148 次
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 被引用 93 次
- Enhanced Regularizers for Attributional RobustnessAnindya Sarkar, Anirban Sarkar, Vineeth N. BalasubramanianAAAI 2021 · 被引用 18 次
相关 Paper
- SAM: The Sensitivity of Attribution Methods to HyperparametersNaman Bansal, Chirag Agarwal, Anh NguyenCVPR 2020
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel 等NeurIPS 2020 · 被引用 67 次
- Pixel-level Certified Explanations via Randomized SmoothingAlaa Anani, Tobias Lorenz, Mario Fritz, Bernt SchieleICML 2025
- On the Robustness of Removal-Based Feature AttributionsChris Lin, Ian Covert, Su-In LeeNeurIPS 2023 · 被引用 25 次
- Towards More Robust Interpretation via Local Gradient AlignmentSunghwan Joo, Seokhyeon Jeong, Juyeon Heo, Adrian Weller 等AAAI 2023 · 被引用 8 次
