Generalizing Backpropagation for Gradient-Based Interpretability
Kevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt, Ryan Cotterell
摘要
Many popular feature-attribution methods for interpreting deep neural networks rely on computing the gradients of a model’s output with respect to its inputs. While these methods can indicate which input features may be important for the model’s prediction, they reveal little about the inner workings of the model itself. In this paper, we observe that the gradient computation of a model is a special case of a more general formulation using semirings. This observation allows us to generalize the backpropagation algorithm to efficiently compute other interpretable statistics about the gradient graph of a neural network, such as the highest-weighted path and entropy. We implement this generalized algorithm, evaluate it on synthetic datasets to better understand the statistics it computes, and apply it to study BERT’s behavior on the subject–verb number agreement task (SVA). With this method, we (a) validate that the amount of gradient flow through a component of a model reflects its importance to a prediction and (b) for SVA, identify which pathways of the self-attention mechanism are most important.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Successor Heads: Recurring, Interpretable Attention Heads In The WildRhys Gould, Euan Ong, George Ogden, Arthur ConmyICLR 2024 · 被引用 75 次
- The Gradient of Algebraic Model CountingJaron Maene, Luc De RaedtAAAI 2025 · 被引用 1 次
它引用的顶会 Paper6
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsXiaochuang Han, Byron C. Wallace, Yulia TsvetkovACL 2020 · 被引用 91 次
- Probing for the Usage of Grammatical NumberKarim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau 等ACL 2022 · 被引用 72 次
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 被引用 34 次
- Influence Patterns for Explaining Information Flow in BERTKaiji Lu, Zifan Wang, Piotr Mardziel, Anupam DattaNeurIPS 2021 · 被引用 22 次
- Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language ModelsKaiji Lu, Piotr Mardziel, Klas Leino, Matt Fredrikson 等ACL 2020 · 被引用 8 次
相关 Paper
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 被引用 282 次
- Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP ModelsSandipan Sikdar, Parantapa Bhattacharya, Kieran HeeseACL 2021
- Efficient Computation of Higher-Order Subgraph Attribution via Message PassingPing Xiong, Thomas Schnake, Grégoire Montavon, Klaus-Robert Müller 等ICML 2022 · 被引用 15 次
- Neural Response Interpretation Through the Lens of Critical PathwaysAshkan Khakzar, Soroosh Baselizadeh, Saurabh Khanduja, Christian Rupprecht 等CVPR 2021
- MFABA: A More Faithful and Accelerated Boundary-Based Attribution Method for Deep Neural NetworksZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang 等AAAI 2024 · 被引用 16 次
