Vision Language Pre-training by Contrastive Learning with Cross-Modal Similarity Regulation
Chaoya Jiang, Wei Ye, Haiyang Xu, Songfang Huang, Fei Huang, Shikun Zhang
Abstract
In this paper, we reconsider the problem of (partial) false negative samples from the Mutual Information (MI) Maximization perspective, the traditional contrastive loss (like InfoNCE loss) will equally push away the anchor of all positive samples and negative samples regardless of their possible semantic similarities. We theoretically show that InfoNCE loss will not only maximize the MI between the anchor and positive samples but minimize the MI between the anchor and false negative samples even though they share similar semantic which could provide a possible theoretical explanation for the observation of the existence of false negative samples in the cross-modal contrastive learning will decrease the downstream task performance of VLP models. Above analysis motivate us to propose the VLP model with a novel Semantic Awared Contrastive Learning framework named SACL where different negative samples are assigned with different contrastive weights according to the semantic similarity between them and the anchor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7bd5dfef-a574-4176-a443-3eaf88fb94a9Cited by top-tier papers7
- COPA : Efficient Vision-Language Pre-training through Collaborative Object- and Patch-Text AlignmentChaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye et al.ACM MM 2023 · 10 citations
- BUS : Efficient and Effective Vision-language Pre-training with Bottom-Up Patch SummarizationChaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye et al.ICCV 2023 · 9 citations
- Improving the Robustness of Knowledge-Grounded Dialogue via Contrastive LearningJiaan Wang, Jianfeng Qu, Kexin Wang, Zhixu Li et al.AAAI 2024 · 5 citations
- Guiding Cross-Modal Representations with MLLM Priors via Preference AlignmentPengfei Zhao, Rongbo Luan, Wei Zhang, Peng Wu et al.NeurIPS 2025 · 3 citations
- FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language AlignmentMyunsoo Kim, Seong-Woong Shim, Byung-Jun LeeCVPR 2026 · 2 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Contrastive Multimodal Fusion with TupleInfoNCEYunze Liu, Qingnan Fan, Shanghang Zhang, Hao Dong et al.ICCV 2021 · 84 citations
- Semantic-Aware Hard Negative Mining for Medical Vision-Language Contrastive PretrainingYongxin Li, Ying Cheng, Yaning Pan, Wen He et al.ACM MM 2025 · 2 citations
- Rethinking Negative Pairs in Code SearchHaochen Li, Xin Zhou, Anh Tuan Luu, Chunyan MiaoEMNLP 2023 · 6 citations
- SNAPHARD CONTRAST LEARNINGChangpu Meng, Jie Yang, Wanqing Li, Yi GuoICLR 2026
- No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive ModelsHai X. Pham, David T. Hoffmann, Ricardo Guerrero, Brais MartínezCVPR 2026 · 1 citation
