XAI for Transformers: Better Explanations through Conservative Propagation
Ameen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon, Klaus-Robert Müller, Lior Wolf
摘要
Transformers have become an important workhorse of machine learning, with numerous applications. This necessitates the development of reliable methods for increasing their transparency. Multiple interpretability methods, often based on gradient information, have been proposed. We show that the gradient in a Transformer reflects the function only locally, and thus fails to reliably identify the contribution of input features to the prediction. We identify Attention Heads and LayerNorm as the main reasons for such unreliable explanations and propose a more stable way for propagation through these layers. Our proposal, which can be seen as a proper extension of the well-established LRP method to Transformers, is shown both theoretically and empirically to overcome the deficiency of a simple gradient-based approach, and achieves state-of-the-art explanation performance on a broad range of Transformer models and datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for TransformersReduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain 等ICML 2024 · 被引用 113 次
- So3krates: Equivariant attention for interactions on arbitrary length-scales in molecular systemsJ. Thorben Frank, Oliver T. Unke, Klaus-Robert MüllerNeurIPS 2022 · 被引用 92 次
- MambaLRP: Explaining Selective State Space Sequence ModelsFarnoush Rezaei Jafari, Grégoire Montavon, Klaus-Robert Müller, Oliver EberleNeurIPS 2024 · 被引用 44 次
- xMIL: Insightful Explanations for Multiple Instance Learning in HistopathologyJulius Hense, Mina Jamshidi Idaji, Oliver Eberle, Thomas Schnake 等NeurIPS 2024 · 被引用 26 次
- Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language ModelsMarvin Pafla, Kate Larson, Mark HancockCHI 2024 · 被引用 16 次
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- Parameterized Explainer for Graph Neural NetworkDongsheng Luo, Wei Cheng, Dongkuan Xu, Wenchao Yu 等NeurIPS 2020 · 被引用 888 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 被引用 451 次
相关 Paper
- Measuring the Mixing of Contextual Information in the TransformerJavier Ferrando, Gerard I. Gállego, Marta R. Costa-jussàEMNLP 2022 · 被引用 18 次
- Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer ExplainabilityYarden Bakish, Itamar Zimerman, Hila Chefer, Lior WolfNeurIPS 2025 · 被引用 8 次
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityWalid Bousselham, Angie W. Boggust, Sofian Chaybouti, Hendrik Strobelt 等ICCV 2025 · 被引用 47 次
- DePass: Unified Feature Attributing by Simple Decomposed Forward PassXiangyu Hong, Che Jiang, Kai Tian, Biqing Qi 等NeurIPS 2025 · 被引用 4 次
- DecompX: Explaining Transformers Decisions by Propagating Token DecompositionAli Modarressi, Mohsen Fayyaz, Ehsan Aghazadeh, Yadollah Yaghoobzadeh 等ACL 2023 · 被引用 7 次
