LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
Faridoun Mehri, Mahdieh Soleymani Baghshah, Mohammad Taher Pilehvar
Abstract
Why do gradient-based explanations struggle with Transformers, and how can we improve them? We identify gradient flow imbalances in Transformers that violate FullGradcompleteness, a critical property for attribution faithfulness that CNNs naturally possess. To address this issue, we introduce LibraGrad-a theoretically grounded post-hoc approach that corrects gradient imbalances through pruning and scaling of backward paths, without changing the forward pass or adding computational overhead. We evaluate LibraGrad using three metric families: Faithfulness, which quantifies prediction changes under perturbations of the most and least relevant features; Completeness Error, which measures attribution conservation relative to model outputs; and Segmentation AP, which assesses alignment with human perception. Extensive experiments across 8 architectures, 4 model sizes, and 5 datasets show that LibraGrad universally enhances gradient-based methods, outperforming existing white-box methods-including Transformerspecific approaches-across all metrics. We demonstrate superior qualitative results through two complementary evaluations: precise text-prompted region highlighting on CLIP models and accurate class discrimination between co-occurring animals on ImageNet-finetuned models-two settings on which existing methods often struggle. Libra-Grad is effective even on the attention-free MLP-Mixer architecture, indicating potential for extension to other modern architectures. Our code is freely available at https: //nightmachinery.github.io/LibraGrad/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Attention (as Discrete-Time Markov) ChainsYotam Erel, Olaf Dünkel, Rishabh Dabral, Vladislav Golyanik et al.NeurIPS 2025 · 12 citations
- Interpreting vision transformers via residual replacement modelJinyeong Kim, Junhyeok Kim, Yumin Shim, Joohyeok Kim et al.NeurIPS 2025 · 4 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for TransformersReduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain et al.ICML 2024 · 113 citations
- Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum FlowBehrooz Azarkhalili, Maxwell W. LibbrechtACL 2025
- Soft Local Completeness: Rethinking Completeness in XAIZiv Weiss Haddad, Oren Barkan, Yehonatan Elisha, Noam KoenigsteinICCV 2025 · 2 citations
- EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit IdentificationLin Zhang, Wenshuo Dong, Zhuoran Zhang, Shu Yang et al.NeurIPS 2025 · 26 citations
- On the Faithfulness of Vision Transformer ExplanationsJunyi Wu, Weitai Kang, Hao Tang, Yuan Hong et al.CVPR 2024 · 8 citations
