When Explanations Lie: Why Many Modified BP Attributions Fail
Leon Sixt, Maximilian Granz, Tim Landgraf
Abstract
Attribution methods aim to explain a neural network's prediction by highlighting the most relevant image areas. A popular approach is to backpropagate (BP) a custom relevance score using modified rules, rather than the gradient. We analyze an extensive set of modified BP methods: Deep Taylor Decomposition, Layer-wise Relevance Propagation (LRP), Excitation BP, PatternAttribution, DeepLIFT, Deconv, RectGrad, and Guided BP. We find empirically that the explanations of all mentioned methods, except for DeepLIFT, are independent of the parameters of later layers. We provide theoretical insights for this surprising behavior and also analyze why DeepLIFT does not suffer from this limitation. Empirically, we measure how information of later layers is ignored by using our new metric, cosine similarity convergence (CSC). The paper provides a framework to assess the faithfulness of new and existing modified BP methods theoretically and empirically. For code see: this https URL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fce971d-6251-4496-a160-98a6491ebd50Cited by top-tier papers33
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
- A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance EstimationThomas Fel, Victor Boutin, Louis Béthune, Rémi Cadène et al.NeurIPS 2023 · 125 citations
- GlanceNets: Interpretable, Leak-proof Concept-based ModelsEmanuele Marconato, Andrea Passerini, Stefano TesoNeurIPS 2022 · 79 citations
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 74 citations
- Rethinking Attention-Model Explainability through Faithfulness Violation TestYibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong et al.ICML 2022 · 60 citations
Builds on3
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
- Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural NetworksWoo-Jeoung Nam, Shir Gur, Jaesik Choi, Lior Wolf et al.AAAI 2020 · 109 citations
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 108 citations
Related papers
- Transformer Interpretability Beyond Attention VisualizationHila Chefer, Shir Gur, Lior WolfCVPR 2021
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 32 citations
- A Unified Taylor Framework for Revisiting Attribution MethodsHuiqi Deng, Na Zou, Mengnan Du, Weifu Chen et al.AAAI 2021 · 25 citations
- Empowering CAM-Based Methods with Capability to Generate Fine-Grained and High-Faithfulness ExplanationsChangqing Qiu, Fusheng Jin, Yining ZhangAAAI 2024 · 11 citations
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon et al.ICML 2022 · 144 citations
