Multi-Granularity Interactive Transformer Hashing for Cross-modal Retrieval
Yishu Liu, Qingpeng Wu, Zheng Zhang, Jingyi Zhang, Guangming Lu
摘要
With the powerful representation ability and privileged efficiency, deep cross-modal hashing (DCMH) has become an emerging fast similarity search technique. Prior studies primarily focus on exploring pairwise similarities across modalities, but fail to comprehensively capture the multi-grained semantic correlations during intra- and inter-modal negotiation. To tackle this issue, this paper proposes a novel Multi-granularity Interactive Transformer Hashing (MITH) network, which hierarchically considers both coarse- and fine-grained similarity measurements across different modalities in one unified transformer-based framework. To the best of our knowledge, this is the first attempt for multi-granularity transformer-based cross-modal hashing. Specifically, a well-designed distilled intra-modal interaction module is deployed to excavate modality-specific concept knowledge with global-local knowledge distillation under the guidance of implicit conceptual category-level representations. Moreover, we construct a contrastive inter-modal alignment module to mine modality-independent semantic concept correspondences with instance- and token-wise contrastive learning, respectively. Such a collaborative learning paradigm can jointly alleviate the heterogeneity and semantic gaps among different modalities from a multi-granularity perspective, yielding discriminative modality-invariant hash codes. Extensive experiments on multiple representative cross-modal datasets demonstrate the consistent superiority of MITH over the existing state-of-the-art baselines. The codes are available at https://github.com/DarrenZZhang/MITH.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 被引用 9 次
- Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken GenerationYongqi Li, Hongru Cai, Wenjie Wang, Leigang Qu 等SIGIR 2025 · 被引用 6 次
- Exploiting Descriptive Completeness Prior for Cross Modal Hashing with Incomplete LabelsHaoyang Luo, Zheng Zhang, Yadan LuoNeurIPS 2024 · 被引用 5 次
- Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text RetrievalYang Du, Yuqi Liu, Qin JinACM MM 2024 · 被引用 4 次
- Asymmetric Cross-Modal Hashing Based on Formal Concept AnalysisYinan Li, Jun Long, Zhan YangAAAI 2025 · 被引用 4 次
相关 Paper
- Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal SearchMeiyu Liang, Junping Du, Zhengyang Liang, Yongwang Xing 等AAAI 2024 · 被引用 24 次
- Bit-aware Semantic Transformer Hashing for Multi-modal RetrievalWentao Tan, Lei Zhu, Weili Guan, Jingjing Li 等SIGIR 2022 · 被引用 33 次
- Stationary and Clustering Transformer Hashing for Cross-modal RetrievalZhan Yang, Yiran Liu, Youyuan Huang, Yinan LiAAAI 2026
- Alleviating the Inconsistency of Multimodal Data in Cross-Modal RetrievalTieying Li, Xiaochun Yang, Yiping Ke, Bin Wang 等ICDE 2024 · 被引用 8 次
- Distribution Consistency Guided Hashing for Cross-Modal RetrievalYuan Sun, Kaiming Liu, Yongxiang Li, Zhenwen Ren 等ACM MM 2024 · 被引用 11 次
