Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers
Stella Frank, Emanuele Bugliarello, Desmond Elliott
摘要
Pretrained vision-and-language BERTs aim to learn representations that combine information from both modalities. We propose a diagnostic method based on cross-modal input ablation to assess the extent to which these models actually integrate cross-modal information. This method involves ablating inputs from one modality, either entirely or selectively based on cross-modal grounding alignments, and evaluating the model prediction performance on the other modality. Model performance is measured by modality-specific tasks that mirror the model pretraining objectives (e.g. masked language modelling for text). Models that have learned to construct cross-modal representations using both modalities are expected to perform worse when inputs are missing from a modality. We find that recently proposed models have much greater relative difficulty predicting text when visual information is ablated, compared to predicting visual object categories when text is ablated, indicating that these models are not symmetrically cross-modal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani 等NeurIPS 2023 · 被引用 462 次
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh 等CVPR 2022 · 被引用 179 次
- When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?Mert Yüksekgönül, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky 等ICLR 2023 · 被引用 37 次
- MultiViz: Towards Visualizing and Understanding Multimodal ModelsPaul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain 等ICLR 2023 · 被引用 15 次
- MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & TasksLetitia Parcalabescu, Anette FrankACL 2023 · 被引用 15 次
它引用的顶会 Paper7
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh 等CVPR 2020
相关 Paper
- MVPTR: Multi-Level Semantic Alignment for Vision-Language Pre-Training via Multi-Stage LearningZejun Li, Zhihao Fan, Huaixiao Tou, Jingjing Chen 等ACM MM 2022 · 被引用 15 次
- DeepAlign: Mitigating Modality Conflict through Modality-Specific AlignmentShuo Li, Bingchen Miao, Wendong Bu, Juncheng Li 等CVPR 2026
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMsYaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan BelinkovNeurIPS 2025 · 被引用 37 次
- BridgeTower: Building Bridges between Encoders in Vision-Language Representation LearningXiao Xu, Chenfei Wu, Shachar Rosenman, Vasudev Lal 等AAAI 2023 · 被引用 99 次
- ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge IntegrationYuhao Cui, Zhou Yu, Chunqi Wang, Zhongzhou Zhao 等ACM MM 2021 · 被引用 46 次
