Influence Patterns for Explaining Information Flow in BERT
Kaiji Lu, Zifan Wang, Piotr Mardziel, Anupam Datta
Abstract
While "attention is all you need" may be proving true, we do not know why: attention-based transformer models such as BERT are superior but how information flows from input tokens to output predictions are unclear. We introduce influence patterns, abstractions of sets of paths through a transformer model. Patterns quantify and localize the flow of information to paths passing through a sequence of model nodes. Experimentally, we find that significant portion of information flow in BERT goes through skip connections instead of attention heads. We further show that consistency of patterns across instances is an indicator of BERT's performance. Finally, we demonstrate that patterns account for far more model performance than previous attention-based and layer-based methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e38904d-45b4-4afa-bcb4-4b3eb248e593Cited by top-tier papers11
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li et al.NeurIPS 2022 · 66 citations
- AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-TuningTao Yang, Jinghao Deng, Xiaojun Quan, Qifan Wang et al.NeurIPS 2022 · 7 citations
- Interpreting Operation Selection in Differentiable Architecture Search: A Perspective from Influence-Directed ExplanationsMiao Zhang, Wei Huang, Bin YangNeurIPS 2022 · 7 citations
- Generalizing Backpropagation for Gradient-Based InterpretabilityKevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt et al.ACL 2023 · 3 citations
- Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMsVishal Pramanik, Maisha Maliha, Nathaniel D. Bastian, Sumit Kumar JhaICLR 2026 · 3 citations
Builds on9
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 799 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter et al.ICLR 2020 · 210 citations
- When BERT Plays the Lottery, All Tickets Are WinningSai Prasanna, Anna Rogers, Anna RumshiskyEMNLP 2020 · 114 citations
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
Related papers
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 282 citations
- Incorporating Residual and Normalization Layers into Analysis of Masked Language ModelsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2021 · 28 citations
- Human Guided Exploitation of Interpretable Attention Patterns in Summarization and Topic SegmentationRaymond Li, Wen Xiao, Linzi Xing, Lanjun Wang et al.EMNLP 2022 · 4 citations
- Block-Skim: Efficient Question Answering for TransformerYue Guan, Zhengyi Li, Zhouhan Lin, Yuhao Zhu et al.AAAI 2022 · 33 citations
- Measuring the Mixing of Contextual Information in the TransformerJavier Ferrando, Gerard I. Gállego, Marta R. Costa-jussàEMNLP 2022 · 18 citations
