AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebastian Lapuschkin, Wojciech Samek
摘要
Large Language Models are prone to biased predictions and hallucinations, underlining the paramount importance of understanding their model-internal reasoning process. However, achieving faithful attributions for the entirety of a black-box transformer model and maintaining computational efficiency is an unsolved challenge. By extending the Layer-wise Relevance Propagation attribution method to handle attention layers, we address these challenges effectively. While partial solutions exist, our method is the first to faithfully and holistically attribute not only input but also latent representations of transformer models with the computational efficiency similar to a single backward pass. Through extensive evaluations against existing methods on LLaMa 2, Mixtral 8x7b, Flan-T5 and vision transformer architectures, we demonstrate that our proposed approach surpasses alternative methods in terms of faithfulness and enables the understanding of latent representations, opening up the door for concept-based explanations. We provide an LRP library at https://github.com/ rachtibat/LRP-eXplains-Transformers .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- MambaLRP: Explaining Selective State Space Sequence ModelsFarnoush Rezaei Jafari, Grégoire Montavon, Klaus-Robert Müller, Oliver EberleNeurIPS 2024 · 被引用 44 次
- The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval AugmentationPatrick Kahardipraja, Reduan Achtibat, Thomas Wiegand, Wojciech Samek 等NeurIPS 2025 · 被引用 13 次
- Attention (as Discrete-Time Markov) ChainsYotam Erel, Olaf Dünkel, Rishabh Dabral, Vladislav Golyanik 等NeurIPS 2025 · 被引用 12 次
- Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer ExplainabilityYarden Bakish, Itamar Zimerman, Hila Chefer, Lior WolfNeurIPS 2025 · 被引用 8 次
- Quantifying Cross-Attention Interaction in Transformers for Interpreting TCR-pMHC BindingJiarui Li, Zixiang Yin, Haley Smith, Zhengming Ding 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 被引用 451 次
相关 Paper
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon 等ICML 2022 · 被引用 144 次
- LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer AttributionsFaridoun Mehri, Mahdieh Soleymani Baghshah, Mohammad Taher PilehvarCVPR 2025
- Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language ModelsZeping Yu, Yonatan Belinkov, Sophia AnaniadouEMNLP 2025 · 被引用 2 次
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceBar Alon, Itamar Zimerman, Lior WolfACL 2026
- DePass: Unified Feature Attributing by Simple Decomposed Forward PassXiangyu Hong, Che Jiang, Kai Tian, Biqing Qi 等NeurIPS 2025 · 被引用 4 次
