Robust Models Are More Interpretable Because Attributions Look Normal
Zifan Wang, Matt Fredrikson, Anupam Datta
Abstract
Recent work has found that adversarially-robust deep networks used for image classification are more interpretable: their feature attributions tend to be sharper, and are more concentrated on the objects associated with the image's ground-truth class. We show that smooth decision boundaries play an important role in this enhanced interpretability, as the model's input gradients around data points will more closely align with boundaries' normal vectors when they are smooth. Thus, because robust models have smoother boundaries, the results of gradient-based attribution methods, like Integrated Gradients and DeepLift, will capture more accurate information about nearby decision boundaries. This understanding of robust interpretability leads to our second contribution: boundary attributions, which aggregate information about the normal vectors of local decision boundaries to explain a classification outcome. We show that by leveraging the key factors underpinning robust interpretability, boundary attributions produce sharper, more concentrated visual explanations -- even on non-robust models. Any example implementation can be found at https://github.com/zifanw/boundary.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Fast Axiomatic Attribution for Neural NetworksRobin Hesse, Simone Schaub-Meyer, Stefan RothNeurIPS 2021 · 55 citations
- On the Relationship Between Explanation and Prediction: A Causal ViewAmir-Hossein Karimi, Krikamol Muandet, Simon Kornblith, Bernhard Schölkopf et al.ICML 2023 · 20 citations
- MFABA: A More Faithful and Accelerated Boundary-Based Attribution Method for Deep Neural NetworksZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang et al.AAAI 2024 · 16 citations
- AttEXplore: Attribution for Explanation with model parameters eXplorationZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang et al.ICLR 2024 · 13 citations
- On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport PerspectiveMathieu Serrurier, Franck Mamalet, Thomas Fel, Louis Béthune et al.NeurIPS 2023 · 11 citations
Builds on10
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Randomized Smoothing of All Shapes and SizesGreg Yang, Tony Duan, J. Edward Hu, Hadi Salman et al.ICML 2020 · 237 citations
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- Globally-Robust Neural NetworksKlas Leino, Zifan Wang, Matt FredriksonICML 2021 · 150 citations
Related papers
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- Hold me tight! Influence of discriminative features on deep network boundariesGuillermo Ortiz-Jiménez, Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2020 · 53 citations
- Towards More Robust Interpretation via Local Gradient AlignmentSunghwan Joo, Seokhyeon Jeong, Juyeon Heo, Adrian Weller et al.AAAI 2023 · 8 citations
- NoiseGrad - Enhancing Explanations by Introducing Stochasticity to Model WeightsKirill Bykov, Anna Hedström, Shinichi Nakajima, Marina M.-C. HöhneAAAI 2022 · 43 citations
