Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification
George Chrysostomou, Nikolaos Aletras
Abstract
Neural network architectures in natural language processing often use attention mechanisms to produce probability distributions over input token representations. Attention has empirically been demonstrated to improve performance in various tasks, while its weights have been extensively used as explanations for model predictions. Recent studies (Jain and Wallace, 2019; Serrano and Smith, 2019; Wiegreffe and Pinter, 2019) have showed that it cannot generally be considered as a faithful explanation (Jacovi and Goldberg, 2020) across encoders and tasks. In this paper, we seek to improve the faithfulness of attention-based explanations for text classification. We achieve this by proposing a new family of Task-Scaling (TaSc) mechanisms that learn task-specific non-contextualised information to scale the original attention weights. Evaluation tests for explanation faithfulness, show that the three proposed variants of TaSc improve attentionbased explanations across two attention mechanisms, five encoders and five text classification datasets without sacrificing predictive performance. Finally, we demonstrate that TaSc consistently provides more faithful attentionbased explanations compared to three widelyused interpretability techniques. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Rethinking Attention-Model Explainability through Faithfulness Violation TestYibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong et al.ICML 2022 · 60 citations
- Flexible Instance-Specific Rationalization of NLP ModelsGeorge Chrysostomou, Nikolaos AletrasAAAI 2022 · 17 citations
- Faithful and Accurate Self-Attention Attribution for Message Passing Neural Networks via the Computation Tree ViewpointYong-Min Shin, Siqing Li, Xin Cao, Won-Yong ShinAAAI 2025 · 6 citations
- Incorporating Attribution Importance for Improving Faithfulness MetricsZhixue Zhao, Nikolaos AletrasACL 2023 · 4 citations
- SPADE: Sparsity-Guided Debugging for Deep Neural NetworksArshia Soltani Moakhar, Eugenia Iofinova, Elias Frantar, Dan AlistarhICML 2024 · 2 citations
Builds on7
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
- Understanding Attention for Text ClassificationXiaobing Sun, Wei LuACL 2020 · 66 citations
- Human Attention Maps for Text Classification: Do Humans and Neural Networks Focus on the Same Words?Cansu Sen, Thomas Hartvigsen, Biao Yin, Xiangnan Kong et al.ACL 2020 · 56 citations
- Learning Variational Word Masks to Improve the Interpretability of Neural Text ClassifiersHanjie Chen, Yangfeng JiEMNLP 2020 · 45 citations
- Learning to Deceive with Attention-Based ExplanationsDanish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig et al.ACL 2020 · 17 citations
Related papers
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li et al.NeurIPS 2022 · 66 citations
- Towards Transparent and Explainable Attention ModelsAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra et al.ACL 2020 · 11 citations
- More Identifiable yet Equally Performant Transformers for Text ClassificationRishabh Bhardwaj, Navonil Majumder, Soujanya Poria, Eduard H. HovyACL 2021
- Attention-based Interpretability with Concept TransformersMattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind et al.ICLR 2022 · 77 citations
- SEAT: Stable and Explainable AttentionLijie Hu, Yixin Liu, Ninghao Liu, Mengdi Huai et al.AAAI 2023 · 30 citations
