Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
Zhijie Nie, Richong Zhang, Zhangchi Feng, Hailang Huang, Xudong Liu
Abstract
Cross-lingual Cross-modal Retrieval (CCR) is an essential task in web search, which aims to break the barriers between modality and language simultaneously and achieves image-text retrieval in the multi-lingual scenario with a single model. In recent years, excellent progress has been made based on cross-lingual cross-modal pre-training; particularly, the methods based on contrastive learning on large-scale data have significantly improved retrieval tasks. However, these methods directly follow the existing pre-training methods in the cross-lingual or cross-modal domain, leading to two problems of inconsistency in CCR: The methods with cross-lingual style suffer from the intra-modal error propagation, resulting in inconsistent recall performance across languages in the whole dataset. The methods with cross-modal style suffer from the inter-modal optimization direction bias, resulting in inconsistent rank across languages within each instance, which cannot be reflected by Recall@K. To solve these problems, we propose a simple but effective 1-to-K contrastive learning method, which treats each language equally and eliminates error propagation and optimization bias. In addition, we propose a new evaluation metric, Mean Rank Variance (MRV), to reflect the rank inconsistency across languages within each instance. Extensive experiments on four CCR datasets show that our method improves both recall rates and MRV with smaller-scale pre-trained data, achieving the new state-of-art 1 . CCS Concepts • Information systems → Image search; multi-lingual and cross-lingual retrieval; Retrieval effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a36c267-271c-4a3b-a263-2ec23dfa49faCited by top-tier papers1
Ask how each one uses itBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsHangbo Bao, Wenhui Wang, Li Dong, Qiang Liu et al.NeurIPS 2022 · 790 citations
Related papers
- Alleviating the Inconsistency of Multimodal Data in Cross-Modal RetrievalTieying Li, Xiaochun Yang, Yiping Ke, Bin Wang et al.ICDE 2024 · 8 citations
- Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-trainingYan Zeng, Wangchunshu Zhou, Ao Luo, Ziming Cheng et al.ACL 2023 · 18 citations
- COOKIE: Contrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language RepresentationKeyu Wen, Jin Xia, Yuanyuan Huang, Linyang Li et al.ICCV 2021 · 35 citations
- Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality InversionMarco Mistretta, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini et al.ICLR 2025
- CAliC: Accurate and Efficient Image-Text Retrieval via Contrastive Alignment and Visual Contexts ModelingHongyu Gao, Chao Zhu, Mengyin Liu, Weibo Gu et al.ACM MM 2022 · 8 citations
