Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching
Xiang Ma, Xuemei Li, Lexin Fang, Caiming Zhang
Abstract
Many contrastive learning based models have achieved advanced performance in image-text matching tasks. The key of these models lies in analyzing the correlation between image-text pairs, which involves cross-modal interaction of embeddings in corresponding dimensions. However, the embeddings of different modalities are from different models or modules, and there is a significant modality gap. Directly interacting such embeddings lacks rationality and may capture inaccurate correlation. Therefore, we propose a novel method called DIAS to bridge the modality gap from two aspects: (1) We align the information representation of embeddings from different modalities in corresponding dimension to ensure the correlation calculation is based on interactions of similar information. (2) The spatial constraints of inter- and intra-modalities unmatched pairs are introduced to ensure the effectiveness of semantic alignment of the model. Besides, a sparse correlation algorithm is proposed to select strong correlated spatial relationships, enabling the model to learn more significant features and avoid being misled by weak correlation. Extensive experiments demonstrate the superiority of DIAS, achieving 4.3%-10.2% rSum improvements on Flickr30k and MSCOCO benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5d475d2-b7a4-4e28-9c35-919e03543445Cited by top-tier papers6
- FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal DataYiting Li, Fayao Liu, Jingyi Liao, Sichao Tian et al.ICCV 2025 · 5 citations
- Guiding Cross-Modal Representations with MLLM Priors via Preference AlignmentPengfei Zhao, Rongbo Luan, Wei Zhang, Peng Wu et al.NeurIPS 2025 · 3 citations
- Reliable Cross-modal Alignment via Prototype Iterative ConstructionXiang Ma, Litian Xu, Lexin Fang, Caiming Zhang et al.ACM MM 2025 · 1 citation
- Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal RepresentationsJeonghyeon Kim, Sangheum HwangCVPR 2025
- Aligning the True Semantics: Constrained Decoupling and Distribution Sampling for Cross-Modal AlignmentXiang Ma, Lexin Fang, Litian Xu, Caiming ZhangAAAI 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- CyCLIP: Cyclic Contrastive Language-Image PretrainingShashank Goel, Hritik Bansal, Sumit Bhatia, Ryan A. Rossi et al.NeurIPS 2022 · 192 citations
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao et al.SIGIR 2021 · 187 citations
Related papers
- Multimodal Aligned Semantic Knowledge for Unpaired Image-text MatchingLaiguo Yin, Yixin Zhang, YUQING SUN, Lizhen CuiICLR 2026
- Learning Relation Alignment for Calibrated Cross-modal RetrievalShuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men et al.ACL 2021
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 12 citations
- Unlocking the Power of Cross-Dimensional Semantic Dependency for Image-Text MatchingKun Zhang, Lei Zhang, Bo Hu, Mengxiao Zhu et al.ACM MM 2023 · 19 citations
- The Convergent Representation of Contrastive Vision-Language Models: Geometry, Modality Gap and Shared Space AlignmentLingjie Yi, Raphael Douady, Chao ChenICML 2026
