Semi-supervised Prototype Semantic Association Learning for Robust Cross-modal Retrieval
Junsheng Wang, Tiantian Gong, Yan Yan
Abstract
Semi-supervised cross-modal retrieval (SS-CMR) aims at learning modality invariance and semantic discrimination from labeled data and unlabeled data, which is crucial for practical applications in the real-world. The key to essentially addressing the SS-CMR task is to solve the semantic association and modality heterogeneity problems. To address these issues, in this paper, we propose a novel semi-supervised cross-modal retrieval method, namely Semi-supervised Prototype Semantic Association Learning (SPAL) for robust cross-modal retrieval. To be specific, we employ shared semantic prototypes to associate labeled and unlabeled data over both modalities to minimize intra-class and maximize inter-class variations, thereby improving discriminative representations on unlabeled data. What is more important is that we propose a novel pseudo-label guided contrastive learning to refine cross-modal representation consistency in the common space, which leverages pseudo-label semantic graph information to constrain cross-modal consistent representations. Meanwhile, multi-modal data inevitably suffers from the cost and difficulty of data collection, resulting in the incomplete multimodal data problem. Thus, to strengthen the robustness of the SS-CMR, we propose a novel prototype propagation method for incomplete data to reconstruct completion representations which preserves the semantic consistency. Extensive evaluations using several baseline methods across four benchmark datasets demonstrate the effectiveness of our method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 77940574-7afe-49e5-9e99-9f40e7ffa193Cited by top-tier papers5
- A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-IdentificationYunpeng Gong, Yongjie Hou, Jiangming Shi, Kim Long Diep et al.AAAI 2026 · 7 citations
- DePro: Domain Ensemble using Decoupled Prompts for Universal Cross-Domain RetrievalKaixiang Chen, Pengfei Fang, Hui XueSIGIR 2025 · 2 citations
- Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person SearchJiahao Zhang, Shaofei Huang, Yaxiong Wang, Zhedong ZhengSIGIR 2026
- Causality-Aligned Semantic Recovery for Incomplete Cross-Modal RetrievalHaipeng Chen, Yu Liu, Xun Yang, Yuheng Liang et al.AAAI 2026
- Robust Semi-paired Multimodal Learning for Cross-modal RetrievalYang Qin, Yuan Sun, Xi Peng, Dezhong Peng et al.AAAI 2026
Related papers
- C3CMR: Cross-Modality Cross-Instance Contrastive Learning for Cross-Media RetrievalJunsheng Wang, Tiantian Gong, Zhixiong Zeng, Changchang Sun et al.ACM MM 2022 · 12 citations
- PAN: Prototype-based Adaptive Network for Robust Cross-modal RetrievalZhixiong Zeng, Shuai Wang, Nan Xu, Wenji MaoSIGIR 2021 · 38 citations
- Partially Aligned Cross-modal Retrieval via Optimal Transport-based Prototype Alignment LearningJunsheng Wang, Tiantian Gong, Yan YanACM MM 2024 · 3 citations
- DiCA: Disambiguated Contrastive Alignment for Cross-Modal Retrieval with Partial LabelsChao Su, Huiming Zheng, Dezhong Peng, Xu WangAAAI 2025 · 8 citations
- Noise Self-Correction via Relation Propagation for Robust Cross-Modal RetrievalRuoxuan Li, Xiangyu Wu, Yang YangACM MM 2025
