Towards Better Understanding Attribution Methods
Sukrut Rao, Moritz Böhle, Bernt Schiele
摘要
Deep neural networks are very successful on many vision tasks, but hard to interpret due to their black box nature. To overcome this, various post-hoc attribution methods have been proposed to identify image regions most influential to the models' decisions. Evaluating such methods is challenging since no ground truth attributions exist. We thus propose three novel evaluation schemes to more reliably measure the faithfulness of those methods, to make comparisons between them more fair, and to make visual inspection more systematic. To address faithfulness, we propose a novel evaluation setting (DiFull) in which we carefully control which parts of the input can influence the output in order to distinguish possible from impossible attributions. To address fairness, we note that different methods are applied at different layers, which skews any comparison, and so evaluate all methods on the same layers (ML-Att) and discuss how this impacts their performance on quantitative metrics. For more systematic visualizations, we propose a scheme (AggAtt) to qualitatively evaluate the methods on complete datasets. We use these evaluation schemes to study strengths and shortcomings of some widely used attribution methods. Finally, we propose a post-processing smoothing step that significantly improves the performance of some attribution methods, and discuss its applicability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance EstimationThomas Fel, Victor Boutin, Louis Béthune, Rémi Cadène 等NeurIPS 2023 · 被引用 125 次
- FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI MethodsRobin Hesse, Simone Schaub-Meyer, Stefan RothICCV 2023 · 被引用 50 次
- Don't trust your eyes: on the (un)reliability of feature visualizationsRobert Geirhos, Roland S. Zimmermann, Blair L. Bilodeau, Wieland Brendel 等ICML 2024 · 被引用 38 次
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 被引用 22 次
- Towards Faithful XAI Evaluation via Generalization-Limited Backdoor WatermarkMengxi Ya, Yiming Li, Tao Dai, Bin Wang 等ICLR 2024 · 被引用 19 次
它引用的顶会 Paper4
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 被引用 220 次
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 被引用 74 次
- There and Back Again: Revisiting Backpropagation Saliency MethodsSylvestre-Alvise Rebuffi, Ruth Fong, Xu Ji, Andrea VedaldiCVPR 2020
- Convolutional Dynamic Alignment Networks for Interpretable ClassificationsMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2021
相关 Paper
- Logic Traps in Evaluating Attribution ScoresYiming Ju, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang 等ACL 2022
- Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions ExplainabilityJoakim Edin, Andreas Geert Motzfeldt, Casper L. Christensen, Tuukka Ruotsalo 等ACL 2025
- Towards credible visual model interpretation with path attributionNaveed Akhtar, Mohammad A. A. K. JalwanaICML 2023 · 被引用 6 次
- Faithfulness Under the Distribution: A New Look at Attribution EvaluationZhiyu Zhu, Zhibo Jin, Jiayu Zhang, Bartlomiej Sobieski 等ICLR 2026
- Towards Better Visualizing the Decision Basis of Networks via Unfold and Conquer Attribution GuidanceJung-Ho Hong, Woo-Jeoung Nam, Kyu-Sung Jeon, Seong-Whan LeeAAAI 2023 · 被引用 3 次
