MASS: Overcoming Language Bias in Image-Text Matching
Jiwan Chung, Seungwon Lim, Sangkyu Lee, Youngjae Yu
Abstract
Pretrained visual-language models have made significant advancements in multimodal tasks, including image-text retrieval. However, a major challenge in image-text matching lies in language bias, where models predominantly rely on language priors and neglect to adequately consider the visual content. We thus present Multimodal ASsociation Score (MASS), a framework that reduces the reliance on language priors for better visual accuracy in image-text matching problems. It can be seamlessly incorporated into existing visual-language models without necessitating additional training. Our experiments have shown that effectively lessens language bias without losing an understanding of linguistic compositionality. Overall, MASS offers a promising solution for enhancing image-text matching performance in visual-language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Revisiting the Role of Language Priors in Vision-Language ModelsZhiqiu Lin, Xinyue Chen, Deepak Pathak, Pengchuan Zhang et al.ICML 2024 · 44 citations
- VL-Match: Enhancing Vision-Language Pretraining with Token-Level and Instance-Level MatchingJunyu Bi, Daixuan Cheng, Ping Yao, Bochen Pang et al.ICCV 2023 · 6 citations
- Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image GuidanceHaozhe Zhao, Shuzheng Si, Liang Chen, Yichi Zhang et al.EMNLP 2025 · 1 citation
- Visually-Augmented Language ModelingWeizhi Wang, Li Dong, Hao Cheng, Haoyu Song et al.ICLR 2023 · 5 citations
- Learning Relation Alignment for Calibrated Cross-modal RetrievalShuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men et al.ACL 2021
