Do Feature Attribution Methods Correctly Attribute Features?
Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie Shah
Abstract
Feature attribution methods are popular in interpretable machine learning. These methods compute the attribution of each input feature to represent its importance, but there is no consensus on the definition of "attribution", leading to many competing methods with little systematic evaluation, complicated in particular by the lack of ground truth attribution. To address this, we propose a dataset modification procedure to induce such ground truth. Using this procedure, we evaluate three common methods: saliency maps, rationales, and attentions. We identify several deficiencies and add new perspectives to the growing body of evidence questioning the correctness and reliability of these methods applied on datasets in the wild. We further discuss possible avenues for remedy and recommend new attribution methods to be tested against ground truth before deployment. The code and appendix are available at https://yilunzhou.github.io/feature-attribution-evaluation/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Tracr: Compiled Transformers as a Laboratory for InterpretabilityDavid Lindner, János Kramár, Sebastian Farquhar, Matthew Rahtz et al.NeurIPS 2023 · 113 citations
- A Comprehensive Study of Image Classification Model Sensitivity to Foregrounds, Backgrounds, and Visual AttributesMazda Moayeri, Phillip Pope, Yogesh Balaji, Soheil FeiziCVPR 2022 · 42 citations
- "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text ClassificationJasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm et al.EMNLP 2022 · 29 citations
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Sanity Simulations for Saliency MethodsJoon Sik Kim, Gregory Plumb, Ameet TalwalkarICML 2022 · 24 citations
Builds on6
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 476 citations
- Benchmarking Deep Learning Interpretability in Time Series PredictionsAya Abdelsalam Ismail, Mohamed K. Gunady, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2020 · 249 citations
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- Learning to Deceive with Attention-Based ExplanationsDanish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig et al.ACL 2020 · 17 citations
Related papers
- Towards Rigorous Interpretations: a Formalisation of Feature AttributionDarius Afchar, Vincent Guigue, Romain HennequinICML 2021 · 22 citations
- A Unified Taylor Framework for Revisiting Attribution MethodsHuiqi Deng, Na Zou, Mengnan Du, Weifu Chen et al.AAAI 2021 · 25 citations
- Evaluating Attribution for Graph Neural NetworksBenjamín Sánchez-Lengeling, Jennifer N. Wei, Brian K. Lee, Emily Reif et al.NeurIPS 2020 · 159 citations
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 32 citations
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 51 citations
