CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
Yabing Wang, Fan Wang, Jianfeng Dong, Hao Luo
Abstract
Cross-lingual cross-modal retrieval has garnered increasing attention recently, which aims to achieve the alignment between vision and target language (V-T) without using any annotated V-T data pairs. Current methods employ machine translation (MT) to construct pseudo-parallel data pairs, which are then used to learn a multi-lingual and multi-modal embedding space that aligns visual and target-language representations. However, the large heterogeneous gap between vision and text, along with the noise present in target language translations, poses significant challenges in effectively aligning their representations. To address these challenges, we propose a general framework, Cross-Lingual to Cross-Modal (CL2CM), which improves the alignment between vision and target language using cross-lingual transfer. This approach allows us to fully leverage the merits of multi-lingual pre-trained models (e.g., mBERT) and the benefits of the same modality structure, i.e., smaller gap, to provide reliable and comprehensive semantic correspondence (knowledge) for the cross-modal network. We evaluate our proposed approach on two multilingual image-text datasets, Multi30K and MSCOCO, and one video-text dataset, VATEX. The results clearly demonstrate the effectiveness of our proposed method and its high potential for large-scale retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 808eacf8-e91b-4024-a4d3-4098e2f05391Cited by top-tier papers9
- Multimodal LLM Enhanced Cross-lingual Cross-modal RetrievalYabing Wang, Le Wang, Qiang Zhou, Zhibin Wang et al.ACM MM 2024 · 24 citations
- Referencing Where to Focus: Improving Visual Grounding with Referential QueryYabing Wang, Zhuotao Tian, Qingpei Guo, Zheng Qin et al.NeurIPS 2024 · 9 citations
- SaCo Loss: Sample-Wise Affinity Consistency for Vision-Language Pre-TrainingSitong Wu, Haoru Tan, Zhuotao Tian, Yukang Chen et al.CVPR 2024 · 5 citations
- CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian SplattingSiyu Jiao, Haoye Dong, Yuyang Yin, Zequn Jie et al.ICCV 2025 · 4 citations
- Aligning Composed Query with Image via Discriminative Perception from Negative CorrespondencesYifan Wang, Wuliang Huang, Chun YuanAAAI 2025 · 3 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Learning with Noisy Correspondence for Cross-modal MatchingZhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding et al.NeurIPS 2021 · 215 citations
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video RetrievalXun Yang, Jianfeng Dong, Yixin Cao, Xun Wang et al.SIGIR 2020 · 131 citations
Related papers
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang et al.ACM MM 2022 · 26 citations
- Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-trainingYan Zeng, Wangchunshu Zhou, Ao Luo, Ziming Cheng et al.ACL 2023 · 18 citations
- CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine TranslationDevaansh Gupta, Siddhant Kharbanda, Jiawei Zhou, Wanhua Li et al.ICCV 2023 · 28 citations
- UC2: Universal Cross-Lingual Cross-Modal Vision-and-Language Pre-TrainingMingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng et al.CVPR 2021
- LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine TranslationHongcheng Guo, Jiaheng Liu, Haoyang Huang, Jian Yang et al.EMNLP 2022 · 9 citations
