MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
Letitia Parcalabescu, Anette Frank
摘要
Vision and language models (VL) are known to exploit unrobust indicators in individual modalities (e.g., introduced by distributional biases) instead of focusing on relevant information in each modality. That a unimodal model achieves similar accuracy on a VL task to a multimodal one, indicates that so-called unimodal collapse occurred. However, accuracybased tests fail to detect e.g., when the model prediction is wrong, while the model used relevant information from a modality. Instead, we propose MM-SHAP, a performance-agnostic multimodality score based on Shapley values that reliably quantifies in which proportions a multimodal model uses individual modalities. We apply MM-SHAP in two ways: (1) to compare models for their average degree of multimodality, and (2) to measure for individual models the contribution of individual modalities for different tasks and datasets. Experiments with six VL models -LXMERT, CLIP and four ALBEF variants -on four VL tasks highlight that unimodal collapse can occur to different degrees and in different directions, contradicting the wide-spread assumption that unimodal collapse is one-sided. Based on our results, we recommend MM-SHAP for analysing multimodal tasks, to diagnose and guide progress towards multimodal integration. Code available at https://github.com/ Heidelberg-NLP/MM-SHAP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Measuring Cross-Modal Interactions in Multimodal ModelsLaura Wenderoth, Konstantin Hemker, Nikola Simidjievski, Mateja JamnikAAAI 2025 · 被引用 12 次
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli 等NeurIPS 2025 · 被引用 11 次
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf InteractionsHubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer 等NeurIPS 2025 · 被引用 6 次
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalDivyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit ChopraICLR 2026 · 被引用 3 次
- SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative RecommendationWei Chen, Xingyu Guo, Shuang Li, Fuwei Zhang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 被引用 476 次
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh 等CVPR 2022 · 被引用 179 次
相关 Paper
- Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal TransformersStella Frank, Emanuele Bugliarello, Desmond ElliottEMNLP 2021 · 被引用 36 次
- Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)Subba Reddy Oota, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa 等ICLR 2025
- Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language ModelsSimon Schrodi, David T. Hoffmann, Max Argus, Volker Fischer 等ICLR 2025
- Ranked from Within: Ranking Large Multimodal Models Without LabelsWeijie Tu, Weijian Deng, Dylan Campbell, Yu Yao 等ICML 2025
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMsHaoran Lou, Chunxiao Fan, Ziyan Liu, Yuexin Wu 等ICCV 2025 · 被引用 1 次
