When Explanations Lie: Why Many Modified BP Attributions Fail
Leon Sixt, Maximilian Granz, Tim Landgraf
摘要
Attribution methods aim to explain a neural network's prediction by highlighting the most relevant image areas. A popular approach is to backpropagate (BP) a custom relevance score using modified rules, rather than the gradient. We analyze an extensive set of modified BP methods: Deep Taylor Decomposition, Layer-wise Relevance Propagation (LRP), Excitation BP, PatternAttribution, DeepLIFT, Deconv, RectGrad, and Guided BP. We find empirically that the explanations of all mentioned methods, except for DeepLIFT, are independent of the parameters of later layers. We provide theoretical insights for this surprising behavior and also analyze why DeepLIFT does not suffer from this limitation. Empirically, we measure how information of later layers is ignored by using our new metric, cosine similarity convergence (CSC). The paper provides a framework to assess the faithfulness of new and existing modified BP methods theoretically and empirically. For code see: this https URL
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 被引用 220 次
- A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance EstimationThomas Fel, Victor Boutin, Louis Béthune, Rémi Cadène 等NeurIPS 2023 · 被引用 125 次
- GlanceNets: Interpretable, Leak-proof Concept-based ModelsEmanuele Marconato, Andrea Passerini, Stefano TesoNeurIPS 2022 · 被引用 79 次
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 被引用 74 次
- Rethinking Attention-Model Explainability through Faithfulness Violation TestYibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong 等ICML 2022 · 被引用 60 次
它引用的顶会 Paper3
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 被引用 220 次
- Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural NetworksWoo-Jeoung Nam, Shir Gur, Jaesik Choi, Lior Wolf 等AAAI 2020 · 被引用 109 次
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 被引用 108 次
相关 Paper
- Transformer Interpretability Beyond Attention VisualizationHila Chefer, Shir Gur, Lior WolfCVPR 2021
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 被引用 32 次
- A Unified Taylor Framework for Revisiting Attribution MethodsHuiqi Deng, Na Zou, Mengnan Du, Weifu Chen 等AAAI 2021 · 被引用 25 次
- Empowering CAM-Based Methods with Capability to Generate Fine-Grained and High-Faithfulness ExplanationsChangqing Qiu, Fusheng Jin, Yining ZhangAAAI 2024 · 被引用 11 次
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon 等ICML 2022 · 被引用 144 次
