Towards Better Understanding Attribution Methods
Sukrut Rao, Moritz Böhle, Bernt Schiele
Abstract
Deep neural networks are very successful on many vision tasks, but hard to interpret due to their black box nature. To overcome this, various post-hoc attribution methods have been proposed to identify image regions most influential to the models' decisions. Evaluating such methods is challenging since no ground truth attributions exist. We thus propose three novel evaluation schemes to more reliably measure the faithfulness of those methods, to make comparisons between them more fair, and to make visual inspection more systematic. To address faithfulness, we propose a novel evaluation setting (DiFull) in which we carefully control which parts of the input can influence the output in order to distinguish possible from impossible attributions. To address fairness, we note that different methods are applied at different layers, which skews any comparison, and so evaluate all methods on the same layers (ML-Att) and discuss how this impacts their performance on quantitative metrics. For more systematic visualizations, we propose a scheme (AggAtt) to qualitatively evaluate the methods on complete datasets. We use these evaluation schemes to study strengths and shortcomings of some widely used attribution methods. Finally, we propose a post-processing smoothing step that significantly improves the performance of some attribution methods, and discuss its applicability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 816852ed-96ec-4fbb-b368-2215709be9d4Cited by top-tier papers14
- A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance EstimationThomas Fel, Victor Boutin, Louis Béthune, Rémi Cadène et al.NeurIPS 2023 · 125 citations
- FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI MethodsRobin Hesse, Simone Schaub-Meyer, Stefan RothICCV 2023 · 50 citations
- Don't trust your eyes: on the (un)reliability of feature visualizationsRobert Geirhos, Roland S. Zimmermann, Blair L. Bilodeau, Wieland Brendel et al.ICML 2024 · 38 citations
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 22 citations
- Towards Faithful XAI Evaluation via Generalization-Limited Backdoor WatermarkMengxi Ya, Yiming Li, Tao Dai, Bin Wang et al.ICLR 2024 · 19 citations
Builds on4
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 74 citations
- There and Back Again: Revisiting Backpropagation Saliency MethodsSylvestre-Alvise Rebuffi, Ruth Fong, Xu Ji, Andrea VedaldiCVPR 2020
- Convolutional Dynamic Alignment Networks for Interpretable ClassificationsMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2021
Related papers
- Logic Traps in Evaluating Attribution ScoresYiming Ju, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang et al.ACL 2022
- Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions ExplainabilityJoakim Edin, Andreas Geert Motzfeldt, Casper L. Christensen, Tuukka Ruotsalo et al.ACL 2025
- Towards credible visual model interpretation with path attributionNaveed Akhtar, Mohammad A. A. K. JalwanaICML 2023 · 6 citations
- Faithfulness Under the Distribution: A New Look at Attribution EvaluationZhiyu Zhu, Zhibo Jin, Jiayu Zhang, Bartlomiej Sobieski et al.ICLR 2026
- Towards Better Visualizing the Decision Basis of Networks via Unfold and Conquer Attribution GuidanceJung-Ho Hong, Woo-Jeoung Nam, Kyu-Sung Jeon, Seong-Whan LeeAAAI 2023 · 3 citations
