Attention Meets Post-hoc Interpretability: A Mathematical Perspective
Gianluigi Lopardo, Frédéric Precioso, Damien Garreau
Abstract
Attention-based architectures, in particular transformers, are at the heart of a technological revolution. Interestingly, in addition to helping obtain state-of-the-art results on a wide range of applications, the attention mechanism intrinsically provides meaningful insights on the internal behavior of the model. Can these insights be used as explanations? Debate rages on. In this paper, we mathematically study a simple attention-based architecture and pinpoint the differences between post-hoc and attention-based explanations. We show that they provide quite different results, and that, despite their limitations, post-hoc methods are capable of capturing more useful insights than merely examining the attention weights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- High-Dimensional Analysis of Single-Layer Attention for Sparse-Token ClassificationNicholas Barnfield, Hugo Cui, Yue M. LuICLR 2026 · 8 citations
- Masks Can Be Distracting: On Context Comprehension in Diffusion Language ModelsJulianna Piskorz, Cristina Pinneri, Alvaro Correia, Motasem Alfarra et al.ICML 2026 · 5 citations
- Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholdsEmanuele Troiani, Hugo Cui, Yatin Dandi, Florent Krzakala et al.ICML 2025
- Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language ModelsYilong Xu, Jinhua Gao, Xiaoming Yu, Yuanhai Xue et al.EMNLP 2025
- Contribution Weights: A Geometrical Analysis of Self-Attention TransformersJake Cunningham, Nicola Muca CironeICML 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter et al.ICLR 2020 · 210 citations
- Thinking Like TransformersGail Weiss, Yoav Goldberg, Eran YahavICML 2021 · 183 citations
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
Related papers
- Token Transformation Matters: Towards Faithful Post-Hoc Explanation for Vision TransformerJunyi Wu, Bin Duan, Weitai Kang, Hao Tang et al.CVPR 2024 · 8 citations
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li et al.NeurIPS 2022 · 66 citations
- Is Attention Explanation? An Introduction to the DebateAdrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens et al.ACL 2022
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 282 citations
- DePass: Unified Feature Attributing by Simple Decomposed Forward PassXiangyu Hong, Che Jiang, Kai Tian, Biqing Qi et al.NeurIPS 2025 · 4 citations
