Robust Models Are More Interpretable Because Attributions Look Normal
Zifan Wang, Matt Fredrikson, Anupam Datta
摘要
Recent work has found that adversarially-robust deep networks used for image classification are more interpretable: their feature attributions tend to be sharper, and are more concentrated on the objects associated with the image's ground-truth class. We show that smooth decision boundaries play an important role in this enhanced interpretability, as the model's input gradients around data points will more closely align with boundaries' normal vectors when they are smooth. Thus, because robust models have smoother boundaries, the results of gradient-based attribution methods, like Integrated Gradients and DeepLift, will capture more accurate information about nearby decision boundaries. This understanding of robust interpretability leads to our second contribution: boundary attributions, which aggregate information about the normal vectors of local decision boundaries to explain a classification outcome. We show that by leveraging the key factors underpinning robust interpretability, boundary attributions produce sharper, more concentrated visual explanations -- even on non-robust models. Any example implementation can be found at https://github.com/zifanw/boundary.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Fast Axiomatic Attribution for Neural NetworksRobin Hesse, Simone Schaub-Meyer, Stefan RothNeurIPS 2021 · 被引用 55 次
- On the Relationship Between Explanation and Prediction: A Causal ViewAmir-Hossein Karimi, Krikamol Muandet, Simon Kornblith, Bernhard Schölkopf 等ICML 2023 · 被引用 20 次
- MFABA: A More Faithful and Accelerated Boundary-Based Attribution Method for Deep Neural NetworksZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang 等AAAI 2024 · 被引用 16 次
- AttEXplore: Attribution for Explanation with model parameters eXplorationZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang 等ICLR 2024 · 被引用 13 次
- On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport PerspectiveMathieu Serrurier, Franck Mamalet, Thomas Fel, Louis Béthune 等NeurIPS 2023 · 被引用 11 次
它引用的顶会 Paper10
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Randomized Smoothing of All Shapes and SizesGreg Yang, Tony Duan, J. Edward Hu, Hadi Salman 等ICML 2020 · 被引用 237 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- Globally-Robust Neural NetworksKlas Leino, Zifan Wang, Matt FredriksonICML 2021 · 被引用 150 次
相关 Paper
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel 等NeurIPS 2020 · 被引用 67 次
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu 等ICML 2020 · 被引用 148 次
- Hold me tight! Influence of discriminative features on deep network boundariesGuillermo Ortiz-Jiménez, Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2020 · 被引用 53 次
- Towards More Robust Interpretation via Local Gradient AlignmentSunghwan Joo, Seokhyeon Jeong, Juyeon Heo, Adrian Weller 等AAAI 2023 · 被引用 8 次
- NoiseGrad - Enhancing Explanations by Introducing Stochasticity to Model WeightsKirill Bykov, Anna Hedström, Shinichi Nakajima, Marina M.-C. HöhneAAAI 2022 · 被引用 43 次
