Gradient-based Visual Explanation for Transformer-based CLIP
Chenyang Zhao, Kun Wang, Xingyu Zeng, Rui Zhao, Antoni B. Chan
Abstract
Significant progress has been achieved on the improvement and downstream usages of the Contrastive Language-Image Pre-training (CLIP) vision-language model, while less attention is paid to the interpretation of CLIP. We propose a Gradient-based visual Explanation method for CLIP (Grad-ECLIP), which interprets the matching result of CLIP for specific input image-text pair. By decomposing the architecture of the encoder and discovering the relationship between the matching similarity and intermediate spatial features, Grad-ECLIP produces effective heat maps that show the influence of image regions or words on the CLIP results. Different from the previous Transformer interpretation methods that focus on the utilization of self-attention maps, which are typically extremely sparse in CLIP, we produce high-quality visual explanations by applying channel and spatial weights on token features. Qualitative and quantitative evaluations verify the superiority of Grad-ECLIP compared with the state-of-the-art methods. A series of analysis are conducted based on our visual explanation results, from which we explore the working mechanism of image-text matching, and the strengths and limitations in attribution identification of CLIP. Codes are available here: https://github.com/Cyang-Zhao/Grad-Eclip .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 649d3fcf-2208-409d-aaea-acba4d1dee44Cited by top-tier papers15
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetShufan Shen, Junshu Sun, Qingming Huang, Shuhui WangNeurIPS 2025 · 13 citations
- Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic SegmentationYajie Liu, Guodong Wang, Jinjin Zhang, Qingjie Liu et al.AAAI 2025 · 3 citations
- PhaseWin Search Framework Enable Efficient Object-Level InterpretationZihan Gu, Ruoyu Chen, Junchi Zhang, Yue Hu et al.CVPR 2026 · 1 citation
- Cross-Domain Attribute Alignment with CLIP: A Rehearsal-Free Approach for Class-Incremental Unsupervised Domain AdaptationKerun Mi, Guoliang Kang, Guangyu Li, Lin Zhao et al.ACM MM 2025 · 1 citation
- Advancing Interpretability of CLIP Representations with Concept Surrogate ModelNhat Hoang-Xuan, Xiyuan Wei, Wanli Xing, Tianbao Yang et al.NeurIPS 2025 · 1 citation
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 451 citations
Related papers
- Class Activation Values: Lucid and Faithful Visual Interpretations for CLIP-based Text-Image RetrievalsPengxu Chen, Huazhong Liu, Jihong Ding, Xinghao Huang et al.SIGIR 2025 · 5 citations
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
- Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional UnderstandingWei Li, Zhen Huang, Xinmei Tian, Le Lu et al.EMNLP 2024
- SpaceCLIP: A Vision-Language Pretraining Framework With Spatial Reconstruction On TextBo Zou, Chao Yang, Chengbin Quan, Youjian ZhaoACM MM 2023 · 1 citation
- FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained UnderstandingDong Jing, Xiaolong He, Yutian Luo, Nanyi Fei et al.NeurIPS 2024 · 70 citations
