"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification
Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, Katja Filippova
摘要
Feature attribution a.k.a. input salience methods which assign an importance score to a feature are abundant but may produce surprisingly different results for the same model on the same input. While differences are expected if disparate definitions of importance are assumed, most methods claim to provide faithful attributions and point at the features most relevant for a model's prediction. Existing work on faithfulness evaluation is not conclusive and does not provide a clear answer as to how different methods are to be compared. Focusing on text classification and the model debugging scenario, our main contribution is a protocol for faithfulness evaluation that makes use of partially synthetic data to obtain ground truth for feature importance ranking. Following the protocol, we do an in-depth analysis of four standard salience method classes on a range of datasets and lexical shortcuts for BERT and LSTM models. We demonstrate that some of the most popular method configurations provide poor results even for simple shortcuts while a method judged to be too simplistic works remarkably well for BERT. * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Protein Design with Guided Discrete DiffusionNate Gruver, Samuel Stanton, Nathan C. Frey, Tim G. J. Rudner 等NeurIPS 2023 · 被引用 246 次
- Post hoc Explanations may be Ineffective for Detecting Unknown Spurious CorrelationJulius Adebayo, Michael Muelly, Harold Abelson, Been KimICLR 2022 · 被引用 102 次
- Stability Guarantees for Feature Attributions with Multiplicative SmoothingAnton Xue, Rajeev Alur, Eric WongNeurIPS 2023 · 被引用 18 次
- On the Impact of Knowledge Distillation for Model InterpretabilityHyeongrok Han, Siwon Kim, Hyun-Soo Choi, Sungroh YoonICML 2023 · 被引用 13 次
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu 等EMNLP 2023 · 被引用 12 次
它引用的顶会 Paper11
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 被引用 552 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- Do Feature Attribution Methods Correctly Attribute Features?Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie ShahAAAI 2022 · 被引用 167 次
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 被引用 158 次
相关 Paper
- Evaluating Attribution for Graph Neural NetworksBenjamín Sánchez-Lengeling, Jennifer N. Wei, Brian K. Lee, Emily Reif 等NeurIPS 2020 · 被引用 159 次
- Learning to Faithfully Rationalize by ConstructionSarthak Jain, Sarah Wiegreffe, Yuval Pinter, Byron C. WallaceACL 2020
- An Empirical Study on Explanations in Out-of-Domain SettingsGeorge Chrysostomou, Nikolaos AletrasACL 2022
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 被引用 32 次
- Explain, Edit, and Understand: Rethinking User Study Design for Evaluating Model ExplanationsSiddhant Arora, Danish Pruthi, Norman M. Sadeh, William W. Cohen 等AAAI 2022 · 被引用 47 次
