Class Activation Values: Lucid and Faithful Visual Interpretations for CLIP-based Text-Image Retrievals
Pengxu Chen, Huazhong Liu, Jihong Ding, Xinghao Huang, Shaojun Zou, Laurence Tianruo Yang
Abstract
Transformer-based text-image matching model, known as CLIP, has garnered significant attention owing to its exceptional performance in text-image retrieval tasks and downstream applications. However, the interpret-ability of CLIP remains underexplored. Existing interpretation methods for Transformers often struggle with incomplete and unreliable attributions within the image and text modalities, respectively. In this paper, we propose a fine-grained interpretation method, termed Class Activation Values (CAV), to provide lucid and faithful visual explanations for CLIP-based text-image retrievals. Specifically, we systematically perform multi-scale accumulation and fusion of class-specific gradients and activation value features to generate high-definition explanations for the image encoder. Furthermore, we present element-wise gradient-based weights to attribute fine-grained relevance between value features and output similarity within the text encoder. The proposed CAV is capable of simultaneously rendering detailed and credible explanations due to its precise feature attribution. Extensive qualitative and quantitative experiments are conducted on the ImageNet-1k and MS COCO datasets, and the experimental results demonstrate that CAV outperforms state-of-the-art interpretation methods in both faithfulness and localization assessments across image and text modalities.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ddcb19a4-8702-4ed6-b9c8-274be220b2b1Cited by top-tier papers1
Ask how each one uses itRelated papers
- Gradient-based Visual Explanation for Transformer-based CLIPChenyang Zhao, Kun Wang, Xingyu Zeng, Rui Zhao et al.ICML 2024 · 24 citations
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li et al.NeurIPS 2022 · 66 citations
- LG-CAV: Train Any Concept Activation Vector with Language GuidanceQihan Huang, Jie Song, Mengqi Xue, Haofei Zhang et al.NeurIPS 2024 · 12 citations
- Boosting the visual interpretability of CLIP via adversarial fine-tuningShizhan Gong, Haoyu Lei, Qi Dou, Farzan FarniaICLR 2025
- FLAIR: VLM with Fine-grained Language-informed Image RepresentationsRui Xiao, Sanghwan Kim, Mariana-Iuliana Georgescu, Zeynep Akata et al.CVPR 2025
