If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions
Reza Esfandiarpoor, Cristina Menghini, Stephen H. Bach
摘要
Recent works often assume that Vision-Language Model (VLM) representations are based on visual attributes like shape. However, it is unclear to what extent VLMs prioritize this information to represent concepts. We propose Extract and Explore (EX2), a novel approach to characterize textual features that are important for VLMs. EX2 uses reinforcement learning to align a large language model with VLM preferences and generates descriptions that incorporate features that are important for the VLM. Then, we inspect the descriptions to identify features that contribute to VLM representations. Using EX2, we find that spurious descriptions have a major role in VLM representations despite providing no helpful information, e.g., Click to enlarge photo of CONCEPT. More importantly, among informative descriptions, VLMs rely significantly on non-visual attributes like habitat (e.g., North America) to represent visual concepts. Also, our analysis reveals that different VLMs prioritize different attributes in their representations. Overall, we show that VLMs do not simply match images to scene descriptions and that non-visual or even spurious descriptions significantly influence their representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modallyDarina Koishigarina, Arnas Uselis, Seong Joon OhICLR 2026 · 被引用 33 次
- FLOSS: Free Lunch in Open-Vocabulary Semantic SegmentationYasser Benigmim, Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc 等ICCV 2025 · 被引用 6 次
- Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot ModelsKaican Li, Weiyan Xie, Yongxiang Huang, Didan Deng 等NeurIPS 2024 · 被引用 5 次
- Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLMJunyuan Mao, Qiankun Li, Linghao Meng, Zhicheng He 等CVPR 2026 · 被引用 1 次
- Attribute-formed Class-specific Concept Space: Endowing Language Bottleneck Model with Better Interpretability and ScalabilityJianyang Zhang, Qianli Luo, Guowu Yang, Wenjing Yang 等CVPR 2025
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Conditional Representation Learning for Customized TasksHonglin Liu, Chao Sun, Peng Hu, Yunfan Li 等NeurIPS 2025 · 被引用 6 次
- RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference DataChenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu 等AAAI 2025
- Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement LearningHaonan Jia, Shichao Dong, Xin Dong, Zenghui Sun 等CVPR 2026
- Verbalized Representation Learning for Interpretable Few-Shot GeneralizationCheng-Fu Yang, Da Yin, Wenbo Hu, Heng Ji 等ICCV 2025 · 被引用 1 次
- A Reinforcement Learning-Based Automatic Video Editing Method Using Pre-trained Vision-Language ModelPanwen Hu, Nan Xiao, Feifei Li, Yongquan Chen 等ACM MM 2023 · 被引用 8 次
