Is Attention Explanation? An Introduction to the Debate
Adrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens, Xiaoou Wang, Thomas François, Patrick Watrin
Abstract
The performance of deep learning models in NLP and other fields of machine learning has led to a rise in their popularity, and so the need for explanations of these models becomes paramount. Attention has been seen as a solution to increase performance, while providing some explanations. However, a debate has started to cast doubt on the explanatory power of attention in neural networks. Although the debate has created a vast literature thanks to contributions from various areas, the lack of communication is becoming more and more tangible. In this paper, we provide a clear overview of the insights on the debate by critically confronting works from these different areas. This holistic vision can be of great interest for future works in all the communities concerned by this debate. We sum up the main challenges spotted in these areas, and we conclude by discussing the most promising future avenues on attention as an explanation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt et al.ICLR 2024 · 119 citations
- Interpretable Image Classification with Adaptive Prototype-based Vision TransformersChiyu Ma, Jon Donnelly, Wenjun Liu, Soroush Vosoughi et al.NeurIPS 2024 · 48 citations
- A Simple Interpretable Transformer for Fine-Grained Image Classification and AnalysisDipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang et al.ICLR 2024 · 27 citations
- Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language ModelsZhisong Zhang, Yan Wang, Xinting Huang, Tianqing Fang et al.ACL 2025 · 22 citations
- Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-JudgeQiyuan Zhang, Yufei Wang, Yuxin Jiang, Liangyou Li et al.ACL 2025 · 16 citations
Builds on10
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter et al.ICLR 2020 · 210 citations
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 138 citations
- Understanding Attention for Text ClassificationXiaobing Sun, Wei LuACL 2020 · 66 citations
- Human Attention Maps for Text Classification: Do Humans and Neural Networks Focus on the Same Words?Cansu Sen, Thomas Hartvigsen, Biao Yin, Xiangnan Kong et al.ACL 2020 · 56 citations
- Why Attentions May Not Be Interpretable?Bing Bai, Jian Liang, Guanhua Zhang, Hao Li et al.KDD 2021 · 51 citations
Related papers
- Attention-based Interpretability with Concept TransformersMattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind et al.ICLR 2022 · 77 citations
- Attention Meets Post-hoc Interpretability: A Mathematical PerspectiveGianluigi Lopardo, Frédéric Precioso, Damien GarreauICML 2024 · 16 citations
- Logic Traps in Evaluating Attribution ScoresYiming Ju, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang et al.ACL 2022
- SEAT: Stable and Explainable AttentionLijie Hu, Yixin Liu, Ninghao Liu, Mengdi Huai et al.AAAI 2023 · 30 citations
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 249 citations
