Multi-Granularity Interactive Transformer Hashing for Cross-modal Retrieval
Yishu Liu, Qingpeng Wu, Zheng Zhang, Jingyi Zhang, Guangming Lu
Abstract
With the powerful representation ability and privileged efficiency, deep cross-modal hashing (DCMH) has become an emerging fast similarity search technique. Prior studies primarily focus on exploring pairwise similarities across modalities, but fail to comprehensively capture the multi-grained semantic correlations during intra- and inter-modal negotiation. To tackle this issue, this paper proposes a novel Multi-granularity Interactive Transformer Hashing (MITH) network, which hierarchically considers both coarse- and fine-grained similarity measurements across different modalities in one unified transformer-based framework. To the best of our knowledge, this is the first attempt for multi-granularity transformer-based cross-modal hashing. Specifically, a well-designed distilled intra-modal interaction module is deployed to excavate modality-specific concept knowledge with global-local knowledge distillation under the guidance of implicit conceptual category-level representations. Moreover, we construct a contrastive inter-modal alignment module to mine modality-independent semantic concept correspondences with instance- and token-wise contrastive learning, respectively. Such a collaborative learning paradigm can jointly alleviate the heterogeneity and semantic gaps among different modalities from a multi-granularity perspective, yielding discriminative modality-invariant hash codes. Extensive experiments on multiple representative cross-modal datasets demonstrate the consistent superiority of MITH over the existing state-of-the-art baselines. The codes are available at https://github.com/DarrenZZhang/MITH.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 923d286c-3abe-4cff-8cc2-8415332bfea2Cited by top-tier papers8
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 9 citations
- Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken GenerationYongqi Li, Hongru Cai, Wenjie Wang, Leigang Qu et al.SIGIR 2025 · 6 citations
- Exploiting Descriptive Completeness Prior for Cross Modal Hashing with Incomplete LabelsHaoyang Luo, Zheng Zhang, Yadan LuoNeurIPS 2024 · 5 citations
- Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text RetrievalYang Du, Yuqi Liu, Qin JinACM MM 2024 · 4 citations
- Asymmetric Cross-Modal Hashing Based on Formal Concept AnalysisYinan Li, Jun Long, Zhan YangAAAI 2025 · 4 citations
Related papers
- Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal SearchMeiyu Liang, Junping Du, Zhengyang Liang, Yongwang Xing et al.AAAI 2024 · 24 citations
- Bit-aware Semantic Transformer Hashing for Multi-modal RetrievalWentao Tan, Lei Zhu, Weili Guan, Jingjing Li et al.SIGIR 2022 · 33 citations
- Stationary and Clustering Transformer Hashing for Cross-modal RetrievalZhan Yang, Yiran Liu, Youyuan Huang, Yinan LiAAAI 2026
- Alleviating the Inconsistency of Multimodal Data in Cross-Modal RetrievalTieying Li, Xiaochun Yang, Yiping Ke, Bin Wang et al.ICDE 2024 · 8 citations
- Distribution Consistency Guided Hashing for Cross-Modal RetrievalYuan Sun, Kaiming Liu, Yongxiang Li, Zhenwen Ren et al.ACM MM 2024 · 11 citations
