PAN: Prototype-based Adaptive Network for Robust Cross-modal Retrieval
Zhixiong Zeng, Shuai Wang, Nan Xu, Wenji Mao
Abstract
In practical applications of cross-modal retrieval, test queries of the retrieval system may vary greatly and come from unknown category. Meanwhile, due to the cost and difficulty of data collection as well as other issues, the available data for cross-modal retrieval are often imbalanced over different modalities. In this paper, we address two important issues to increase the robustness of cross-modal retrieval system for real-world applications: handling test queries from unknown category and modality-imbalanced training data. The first issue has not been addressed by existing methods and the second issue was not well addressed in the related research. To tackle the above issues, we take the advantage of prototype learning, and propose a prototype-based adaptive network (PAN) for robust cross-modal retrieval. Our method leverages a unified prototype to represent each semantic category across modalities, which provides discriminative information of different categories and takes unified prototypes as anchors to learn cross-modal representations adaptively. Moreover, we propose a novel prototype propagation strategy to reconstruct balanced representations which preserves the semantic consistency and modality heterogeneity. Experimental results on the benchmark datasets demonstrate the effectiveness of our method compared to the SOTA methods, and further robustness tests show the superiority of our method in solving the above issues.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3fc505af-9af3-4ea8-b7c1-cf2291e57669Cited by top-tier papers3
- Prototype-guided Cross-modal Completion and Alignment for Incomplete Text-based Person Re-identificationTiantian Gong, Guodong Du, Junsheng Wang, Yongkang Ding et al.ACM MM 2023 · 9 citations
- Fine-grained Prototypical Voting with Heterogeneous Mixup for Semi-supervised 2D-3D Cross-modal RetrievalFan Zhang, Xian-Sheng Hua, Chong Chen, Xiao LuoCVPR 2024 · 5 citations
- Causality-Aligned Semantic Recovery for Incomplete Cross-Modal RetrievalHaipeng Chen, Yu Liu, Xun Yang, Yuheng Liang et al.AAAI 2026
Related papers
- Semi-supervised Prototype Semantic Association Learning for Robust Cross-modal RetrievalJunsheng Wang, Tiantian Gong, Yan YanSIGIR 2024 · 3 citations
- MCCN: Multimodal Coordinated Clustering Network for Large-Scale Cross-modal RetrievalZhixiong Zeng, Ying Sun, Wenji MaoACM MM 2021 · 20 citations
- Partially Aligned Cross-modal Retrieval via Optimal Transport-based Prototype Alignment LearningJunsheng Wang, Tiantian Gong, Yan YanACM MM 2024 · 3 citations
- Unbiased Prototype Consistency Learning for Multi-Modal and Multi-Task Object Re-IdentificationZhongao Zhou, Bin Yang, Wenke Huang, Jun Chen et al.NeurIPS 2025 · 2 citations
- DiCA: Disambiguated Contrastive Alignment for Cross-Modal Retrieval with Partial LabelsChao Su, Huiming Zheng, Dezhong Peng, Xu WangAAAI 2025 · 8 citations
