Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective
Emmanuelle Salin, Badreddine Farah, Stéphane Ayache, Benoît Favre
摘要
In recent years, joint text-image embeddings have significantly improved thanks to the development of transformer-based Vision-Language models. Despite these advances, we still need to better understand the representations produced by those models. In this paper, we compare pre-trained and fine-tuned representations at a vision, language and multimodal level. To that end, we use a set of probing tasks to evaluate the performance of state-of-the-art Vision-Language models and introduce new datasets specifically for multimodal probing. These datasets are carefully designed to address a range of multimodal capabilities while minimizing the potential for models to rely on bias. Although the results confirm the ability of Vision-Language models to understand color at a multimodal level, the models seem to prefer relying on bias in text data for object position and size. On semantically adversarial examples, we find that those models are able to pinpoint fine-grained multimodal differences. Finally, we also notice that fine-tuning a Vision-Language model on multimodal tasks does not necessarily improve its multimodal ability. We make all datasets and code available to replicate experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language TransformersEstelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu 等CVPR 2022 · 被引用 34 次
- Grounded Image Text Matching with Mismatched Relation ReasoningYu Wu, Yana Wei, Haozhe Wang, Yongfei Liu 等ICCV 2023 · 被引用 14 次
- Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination MitigationQiming Li, Zekai Ye, Xiaocheng Feng, Weihong Zhong 等AAAI 2026 · 被引用 10 次
- Measuring Progress in Fine-grained Vision-and-Language UnderstandingEmanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks 等ACL 2023 · 被引用 9 次
- Interpretable Debiasing of Vision-Language Models for Social FairnessNa Min An, Yoonna Jang, Yusuke Hirota, Ryo Hachiuma 等CVPR 2026 · 被引用 7 次
它引用的顶会 Paper6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Seeing Out of the Box: End-to-End Pre-Training for Vision-Language Representation LearningZhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu 等CVPR 2021
- Less Is More: ClipBERT for Video-and-Language Learning via Sparse SamplingJie Lei, Linjie Li, Luowei Zhou, Zhe Gan 等CVPR 2021
相关 Paper
- Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive EvaluationHong-Tao Yu, Yuxin Peng, Serge J. Belongie, Xiu-Shen WeiICLR 2026 · 被引用 21 次
- Words or Vision: Do Vision-Language Models Have Blind Faith in Text?Ailin Deng, Tri Cao, Zhirui Chen, Bryan HooiCVPR 2025
- Distribution-Aware Prompt Tuning for Vision-Language ModelsEulrang Cho, Jooyeon Kim, Hyunwoo J. KimICCV 2023 · 被引用 54 次
- Do Vision-and-Language Transformers Learn Grounded Predicate-Noun Dependencies?Mitja Nikolaus, Emmanuelle Salin, Stéphane Ayache, Abdellah Fourtassi 等EMNLP 2022 · 被引用 5 次
- Token Embeddings Alignment for Cross-Modal RetrievalChen-Wei Xie, Jianmin Wu, Yun Zheng, Pan Pan 等ACM MM 2022 · 被引用 18 次
