Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations Interpretability
Zhiyu Zhu, Zhibo Jin, Jiayu Zhang, Nan Yang, Jiahao Huang, Jianlong Zhou, Fang Chen
摘要
The task of identifying multimodal image-text representations has garnered increasing attention, particularly with models such as CLIP (Contrastive Language-Image Pretraining), which demonstrate exceptional performance in learning complex associations between images and text. Despite these advancements, ensuring the interpretability of such models is paramount for their safe deployment in realworld applications, such as healthcare. While numerous interpretability methods have been developed for unimodal tasks, these approaches often fail to transfer effectively to multimodal contexts due to inherent differences in the representation structures. Bottleneck methods, well-established in information theory, have been applied to enhance CLIP's interpretability. However, they are often hindered by strong assumptions or intrinsic randomness. To overcome these challenges, we propose the Narrowing Information Bottleneck Theory, a novel framework that fundamentally redefines the traditional bottleneck approach. This theory is specifically designed to satisfy contemporary attribution axioms, providing a more robust and reliable solution for improving the interpretability of multimodal models. In our experiments, compared to state-of-the-art methods, our approach enhances image interpretability by an average of 9%, text interpretability by an average of 58.83%, and accelerates processing speed by 63.95%. Our code is publicly accessible at https://github.com/LMBTough/NIB .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf InteractionsHubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer 等NeurIPS 2025 · 被引用 6 次
- A Comprehensive Information-Decomposition Analysis of Large Vision-Language ModelsLixin Xiu, Xufang Luo, Hideki NakayamaICLR 2026 · 被引用 4 次
- Advancing Interpretability of CLIP Representations with Concept Surrogate ModelNhat Hoang-Xuan, Xiyuan Wei, Wanli Xing, Tianbao Yang 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 被引用 451 次
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 被引用 220 次
- CLIP-Forge: Towards Zero-Shot Text-to-Shape GenerationAditya Sanghi, Hang Chu, Joseph G. Lambourne, Ye Wang 等CVPR 2022 · 被引用 206 次
相关 Paper
- Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck AttributionYing Wang, Tim G. J. Rudner, Andrew Gordon WilsonNeurIPS 2023 · 被引用 51 次
- V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept TokenizerHangzhou He, Lei Zhu, Xinliang Zhang, Shuang Zeng 等AAAI 2025 · 被引用 11 次
- Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual RepresentationsShizhan Gong, Xiaofan Zhang, Qi DouAAAI 2026
- Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image ClassificationYue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin 等CVPR 2023
- There Was Never a Bottleneck in Concept Bottleneck ModelsAntonio Almudévar, José Miguel Hernández-Lobato, Alfonso OrtegaICLR 2026 · 被引用 9 次
