Dual-Branch Multi-Granularity Network with Structured Contrastive Ranking for Cross-Modal Retrieval
Zihao Chen, Chenyang Bu, Shengwei Ji, Xindong Wu
Abstract
Cross-modal retrieval (CMR) has advanced considerably by mapping image and text features into a shared embedding space; however, these approaches still face two persistent challenges: (1) semantic sparsity, where discriminative cues are confined to localized regions, making it difficult to identify implicit visual evidence; and (2) ranking uncertainty under semantic ambiguity, where models struggle to maintain the correct retrieval order when candidates share similar contexts. To address these issues, we propose the Dual-Branch Multi-Granularity Network (DBMG) with Structured Contrastive Ranking, which enriches visual semantics by leveraging a multimodal large language model to generate auxiliary descriptions, aligns sparse cues through a dual-branch architecture capturing both global and local interactions, and enforces ranking consistency via a three-stage contrastive objective that progressively optimizes category clustering, instance alignment, and margin-based ranking. Extensive experiments on four standard CMR benchmarks demonstrate that DBMG outperforms 12 strong baselines, achieving an average 15.91% improvement in mAP, establishing a new state-of-the-art. The code is available at https://github.com/DMiC-Lab-HFUT/DBMG.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 26c9a5c9-062b-4d01-b2a6-53b42f2a24d6Cited by top-tier papers1
Ask how each one uses itRelated papers
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 12 citations
- SEPS: Semantic-Enhanced Patch Slimming Framework for Fine-Grained Cross-Modal AlignmentXinyu Mao, Junsi Li, Haoji Zhang, Yu Liang et al.ICML 2026
- MCCN: Multimodal Coordinated Clustering Network for Large-Scale Cross-modal RetrievalZhixiong Zeng, Ying Sun, Wenji MaoACM MM 2021 · 20 citations
- Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive LearningZhijie Nie, Richong Zhang, Zhangchi Feng, Hailang Huang et al.KDD 2024 · 4 citations
- C3CMR: Cross-Modality Cross-Instance Contrastive Learning for Cross-Media RetrievalJunsheng Wang, Tiantian Gong, Zhixiong Zeng, Changchang Sun et al.ACM MM 2022 · 12 citations
