A Diagnostic Study of Explainability Techniques for Text Classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein
摘要
Recent developments in machine learning have introduced models that approach human performance at the cost of increased architectural complexity. Efforts to make the rationales behind the models' predictions transparent have inspired an abundance of new explainability techniques. Provided with an already trained model, they compute saliency scores for the words of an input instance. However, there exists no definitive guide on (i) how to choose such a technique given a particular application task and model architecture, and (ii) the benefits and drawbacks of using each such technique. In this paper, we develop a comprehensive list of diagnostic properties for evaluating existing explainability techniques. We then employ the proposed list to compare a set of diverse explainability techniques on downstream text classification tasks and neural network architectures. We also compare the saliency scores assigned by the explainability techniques with human annotations of salient input regions to find relations between a model's performance and the agreement of its rationales with human ones. Overall, we find that the gradient-based explanations perform best across tasks and model architectures, and we present further insights into the properties of the reviewed explainability techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon 等ICML 2022 · 被引用 144 次
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li 等NeurIPS 2022 · 被引用 66 次
- Rethinking Attention-Model Explainability through Faithfulness Violation TestYibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong 等ICML 2022 · 被引用 60 次
- FactKG: Fact Verification via Reasoning on Knowledge GraphsJiho Kim, Sungjin Park, Yeonsu Kwon, Yohan Jo 等ACL 2023 · 被引用 36 次
- On the Sensitivity and Stability of Model Interpretations in NLPFan Yin, Zhouxing Shi, Cho-Jui Hsieh, Kai-Wei ChangACL 2022 · 被引用 35 次
它引用的顶会 Paper3
- Generating Fact Checking ExplanationsPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinACL 2020 · 被引用 130 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- Obtaining Faithful Interpretations from Compositional Neural NetworksSanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson 等ACL 2020 · 被引用 5 次
相关 Paper
- Passive attention in artificial neural networks predicts human visual selectivityThomas A. Langlois, H. Charles Zhao, Erin Grant, Ishita Dasgupta 等NeurIPS 2021 · 被引用 19 次
- Diagnostics-Guided Explanation GenerationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinAAAI 2022 · 被引用 10 次
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 被引用 62 次
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 被引用 51 次
- Scaling Symbolic Methods using Gradients for Neural Model ExplanationSubham Sekhar Sahoo, Subhashini Venugopalan, Li Li, Rishabh Singh 等ICLR 2021 · 被引用 8 次
