Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution
Ying Wang, Tim G. J. Rudner, Andrew Gordon Wilson
摘要
Vision-language pretrained models have seen remarkable success, but their application to safety-critical settings is limited by their lack of interpretability. To improve the interpretability of vision-language models such as CLIP, we propose a multi-modal information bottleneck (M2IB) approach that learns latent representations that compress irrelevant information while preserving relevant visual and textual features. We demonstrate how M2IB can be applied to attribution analysis of vision-language pretrained models, increasing attribution accuracy and improving the interpretability of such models when applied to safety-critical domains such as healthcare. Crucially, unlike commonly used unimodal attribution methods, M2IB does not require ground truth labels, making it possible to audit representations of vision-language pretrained models when multiple modalities but no ground-truth data is available. Using CLIP as an example, we demonstrate the effectiveness of M2IB attribution and show that it outperforms gradient-based, perturbation-based, and attention-based attribution methods both qualitatively and quantitatively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Gradient-based Visual Explanation for Transformer-based CLIPChenyang Zhao, Kun Wang, Xingyu Zeng, Rui Zhao 等ICML 2024 · 被引用 24 次
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 被引用 17 次
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf InteractionsHubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer 等NeurIPS 2025 · 被引用 6 次
- A Comprehensive Information-Decomposition Analysis of Large Vision-Language ModelsLixin Xiu, Xufang Luo, Hideki NakayamaICLR 2026 · 被引用 4 次
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance ApproachAishwarya Agarwal, Srikrishna Karanam, Vineet GandhiCVPR 2026 · 被引用 2 次
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 被引用 451 次
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 被引用 220 次
相关 Paper
- Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations InterpretabilityZhiyu Zhu, Zhibo Jin, Jiayu Zhang, Nan Yang 等ICLR 2025
- V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept TokenizerHangzhou He, Lei Zhu, Xinliang Zhang, Shuang Zeng 等AAAI 2025 · 被引用 11 次
- Boosting the visual interpretability of CLIP via adversarial fine-tuningShizhan Gong, Haoyu Lei, Qi Dou, Farzan FarniaICLR 2025
- Explaining CLIP Zero-shot Predictions Through ConceptsOnat Özdemir, Anders Christensen, Stephan Alaniz, Zeynep Akata 等CVPR 2026 · 被引用 2 次
- Stabilizing Cross-Modal Bidirectional Attribution: Few-Shot Adversarial Prompt Tuning for Robust Vision-Language ModelsJun Feng, Shuhong Wu, Hong Sun, Pengfei Zhang 等AAAI 2026
