C3CMR: Cross-Modality Cross-Instance Contrastive Learning for Cross-Media Retrieval
Junsheng Wang, Tiantian Gong, Zhixiong Zeng, Changchang Sun, Yan Yan
Abstract
Cross-modal retrieval is an essential area of representation learning, which aims to retrieve instances with the same semantics from different modalities. In real implementation, a key challenge for cross-modal retrieval is to narrow the heterogeneity gap between different modalities and obtain modality-invariant and discriminative features. Typically, existing approaches for this task mainly learn inter-modal invariance and focus on how to combine pair-level loss and class-level loss, which cannot effectively and adequately learn discriminative features. To address these issues, in this paper, we propose a novel Cross-Modality Cross-Instance Contrastive Learning for Cross-Media Retrieval (C3CMR) method. Specifically, to fully employ the intra-modal similarities, we introduce the intra-modal contrastive learning to enhance the discriminative power of the unimodal features. Besides, we design a supervised inter-modal contrastive learning scheme to take full advantage of the label semantic associations. In this way, cross-semantic associations and inter-modal invariance can be further learned. Moreover, pertaining to the local suboptimal semantic similarity by only mining pairwise and triplewise sample relationships, we propose the cross-instance contrastive learning to mine the similarities among multiple instances. Comprehensive experimental results on four widely-used benchmark datasets demonstrate the superiority of our proposed method over several state-of-the-art cross-modal retrieval methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Interactive Cross-modal Learning for Text-3D Scene RetrievalYanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin et al.NeurIPS 2025 · 9 citations
- PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape MappingSubash Khanal, Eric Xing, Srikumar Sastry, Aayush Dhakal et al.ACM MM 2024 · 3 citations
- Causality-Aligned Semantic Recovery for Incomplete Cross-Modal RetrievalHaipeng Chen, Yu Liu, Xun Yang, Yuheng Liang et al.AAAI 2026
Related papers
- Semi-supervised Prototype Semantic Association Learning for Robust Cross-modal RetrievalJunsheng Wang, Tiantian Gong, Yan YanSIGIR 2024 · 3 citations
- Alleviating the Inconsistency of Multimodal Data in Cross-Modal RetrievalTieying Li, Xiaochun Yang, Yiping Ke, Bin Wang et al.ICDE 2024 · 8 citations
- MCCN: Multimodal Coordinated Clustering Network for Large-Scale Cross-modal RetrievalZhixiong Zeng, Ying Sun, Wenji MaoACM MM 2021 · 20 citations
- DiCA: Disambiguated Contrastive Alignment for Cross-Modal Retrieval with Partial LabelsChao Su, Huiming Zheng, Dezhong Peng, Xu WangAAAI 2025 · 8 citations
- Cross-Modal Center Loss for 3D Cross-Modal RetrievalLonglong Jing, Elahe Vahdani, Jiaxing Tan, Yingli TianCVPR 2021
