DecompX: Explaining Transformers Decisions by Propagating Token Decomposition
Ali Modarressi, Mohsen Fayyaz, Ehsan Aghazadeh, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar
Abstract
An emerging solution for explaining Transformer-based models is to use vectorbased analysis on how the representations are formed. However, providing a faithful vector-based explanation for a multi-layer model could be challenging in three aspects: (1) Incorporating all components into the analysis, (2) Aggregating the layer dynamics to determine the information flow and mixture throughout the entire model, and (3) Identifying the connection between the vector-based analysis and the model's predictions. In this paper, we present DecompX to tackle these challenges. DecompX is based on the construction of decomposed token representations and their successive propagation throughout the model without mixing them in between layers. Additionally, our proposal provides multiple advantages over existing solutions for its inclusion of all encoder components (especially nonlinear feed-forward networks) and the classification head. The former allows acquiring precise vectors while the latter transforms the decomposition into meaningful prediction-based values, eliminating the need for norm-or summation-based vector aggregation. According to the standard faithfulness evaluations, DecompX consistently outperforms existing gradient-based and vector-based approaches on various datasets. Our code is available at github.com/mohsenfayyaz/DecompX.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f10d914-755d-4beb-bcc5-af1076faa357Cited by top-tier papers14
- Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention MapsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiICLR 2024 · 35 citations
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt et al.ICLR 2026 · 28 citations
- Unveiling and Manipulating Prompt Influence in Large Language ModelsZijian Feng, Hanzhang Zhou, Zixiao Zhu, Junlang Qian et al.ICLR 2024 · 9 citations
- ReX: A Framework for Incorporating Temporal Information in Model-Agnostic Local Explanation TechniquesJunhao Liu, Xin ZhangAAAI 2025 · 6 citations
- An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare RecordsJoakim Edin, Maria Maistro, Lars Maaløe, Lasse Borgholt et al.EMNLP 2024 · 4 citations
Builds on11
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter et al.ICLR 2020 · 210 citations
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 138 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 85 citations
Related papers
- DePass: Unified Feature Attributing by Simple Decomposed Forward PassXiangyu Hong, Che Jiang, Kai Tian, Biqing Qi et al.NeurIPS 2025 · 4 citations
- Measuring the Mixing of Contextual Information in the TransformerJavier Ferrando, Gerard I. Gállego, Marta R. Costa-jussàEMNLP 2022 · 18 citations
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li et al.NeurIPS 2022 · 66 citations
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon et al.ICML 2022 · 144 citations
- Token Transformation Matters: Towards Faithful Post-Hoc Explanation for Vision TransformerJunyi Wu, Bin Duan, Weitai Kang, Hao Tang et al.CVPR 2024 · 8 citations
