Class Activation Values: Lucid and Faithful Visual Interpretations for CLIP-based Text-Image Retrievals
Pengxu Chen, Huazhong Liu, Jihong Ding, Xinghao Huang, Shaojun Zou, Laurence Tianruo Yang
摘要
Transformer-based text-image matching model, known as CLIP, has garnered significant attention owing to its exceptional performance in text-image retrieval tasks and downstream applications. However, the interpret-ability of CLIP remains underexplored. Existing interpretation methods for Transformers often struggle with incomplete and unreliable attributions within the image and text modalities, respectively. In this paper, we propose a fine-grained interpretation method, termed Class Activation Values (CAV), to provide lucid and faithful visual explanations for CLIP-based text-image retrievals. Specifically, we systematically perform multi-scale accumulation and fusion of class-specific gradients and activation value features to generate high-definition explanations for the image encoder. Furthermore, we present element-wise gradient-based weights to attribute fine-grained relevance between value features and output similarity within the text encoder. The proposed CAV is capable of simultaneously rendering detailed and credible explanations due to its precise feature attribution. Extensive qualitative and quantitative experiments are conducted on the ImageNet-1k and MS COCO datasets, and the experimental results demonstrate that CAV outperforms state-of-the-art interpretation methods in both faithfulness and localization assessments across image and text modalities.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Gradient-based Visual Explanation for Transformer-based CLIPChenyang Zhao, Kun Wang, Xingyu Zeng, Rui Zhao 等ICML 2024 · 被引用 24 次
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li 等NeurIPS 2022 · 被引用 66 次
- LG-CAV: Train Any Concept Activation Vector with Language GuidanceQihan Huang, Jie Song, Mengqi Xue, Haofei Zhang 等NeurIPS 2024 · 被引用 12 次
- Boosting the visual interpretability of CLIP via adversarial fine-tuningShizhan Gong, Haoyu Lei, Qi Dou, Farzan FarniaICLR 2025
- FLAIR: VLM with Fine-grained Language-informed Image RepresentationsRui Xiao, Sanghwan Kim, Mariana-Iuliana Georgescu, Zeynep Akata 等CVPR 2025
