A Diagnostic Study of Explainability Techniques for Text Classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein
Abstract
Recent developments in machine learning have introduced models that approach human performance at the cost of increased architectural complexity. Efforts to make the rationales behind the models' predictions transparent have inspired an abundance of new explainability techniques. Provided with an already trained model, they compute saliency scores for the words of an input instance. However, there exists no definitive guide on (i) how to choose such a technique given a particular application task and model architecture, and (ii) the benefits and drawbacks of using each such technique. In this paper, we develop a comprehensive list of diagnostic properties for evaluating existing explainability techniques. We then employ the proposed list to compare a set of diverse explainability techniques on downstream text classification tasks and neural network architectures. We also compare the saliency scores assigned by the explainability techniques with human annotations of salient input regions to find relations between a model's performance and the agreement of its rationales with human ones. Overall, we find that the gradient-based explanations perform best across tasks and model architectures, and we present further insights into the properties of the reviewed explainability techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9edbb5bb-f849-4015-a2f6-8c5c500b937fCited by top-tier papers54
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon et al.ICML 2022 · 144 citations
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li et al.NeurIPS 2022 · 66 citations
- Rethinking Attention-Model Explainability through Faithfulness Violation TestYibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong et al.ICML 2022 · 60 citations
- FactKG: Fact Verification via Reasoning on Knowledge GraphsJiho Kim, Sungjin Park, Yeonsu Kwon, Yohan Jo et al.ACL 2023 · 36 citations
- On the Sensitivity and Stability of Model Interpretations in NLPFan Yin, Zhouxing Shi, Cho-Jui Hsieh, Kai-Wei ChangACL 2022 · 35 citations
Builds on3
- Generating Fact Checking ExplanationsPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinACL 2020 · 130 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Obtaining Faithful Interpretations from Compositional Neural NetworksSanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson et al.ACL 2020 · 5 citations
Related papers
- Passive attention in artificial neural networks predicts human visual selectivityThomas A. Langlois, H. Charles Zhao, Erin Grant, Ishita Dasgupta et al.NeurIPS 2021 · 19 citations
- Diagnostics-Guided Explanation GenerationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinAAAI 2022 · 10 citations
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 62 citations
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 51 citations
- Scaling Symbolic Methods using Gradients for Neural Model ExplanationSubham Sekhar Sahoo, Subhashini Venugopalan, Li Li, Rishabh Singh et al.ICLR 2021 · 8 citations
