Differentiable Cross-modal Hashing via Multimodal Transformers
Junfeng Tu, Xueliang Liu, Zongxiang Lin, Richang Hong, Meng Wang
Abstract
Cross-modal hashing aims at projecting the cross modal content into a common Hamming space for efficient search. Most existing work first encodes the samples with a deep network and then binaries the encoded feature into hashing code. However, the relative location information in the image may be lost when an image is encoded by the convolutional network, which makes it challenging to model the relationship of different modalities. Moreover, it is NP-hard to optimize the model with the discrete sign binary function popularly used in existing solutions. To address these issues, we propose a differentiable cross-modal hashing method that utilizes the multimodal transformer as the backbone to capture the location information in an image when encoding the visual content. In addition, a novel differentiable cross-modal hashing method is proposed to generate the binary code by a selecting mechanism, which could be formulated as a continuous and easily optimized problem. We perform extensive experiments on several cross modal datasets and the results show that the proposed method outperforms many existing solutions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 248b857f-289c-4b0c-898e-d86eb5e1c5caCited by top-tier papers14
- An End-To-End Graph Attention Network Hashing for Cross-Modal RetrievalHuilong Jin, Yingxue Zhang, Lei Shi, Shuang Zhang et al.NeurIPS 2024 · 18 citations
- Pairwise-Label-Based Deep Incremental Hashing with Simultaneous Code ExpansionDayan Wu, Qinghang Su, Bo Li, Weiping WangAAAI 2024 · 11 citations
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen et al.ACM MM 2023 · 11 citations
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 9 citations
- Exploiting Descriptive Completeness Prior for Cross Modal Hashing with Incomplete LabelsHaoyang Luo, Zheng Zhang, Yadan LuoNeurIPS 2024 · 5 citations
Related papers
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 261 citations
- Multi-Granularity Interactive Transformer Hashing for Cross-modal RetrievalYishu Liu, Qingpeng Wu, Zheng Zhang, Jingyi Zhang et al.ACM MM 2023 · 44 citations
- Supervised Hierarchical Deep Hashing for Cross-Modal RetrievalYu-Wei Zhan, Xin Luo, Yongxin Wang, Xin-Shun XuACM MM 2020 · 55 citations
- Stationary and Clustering Transformer Hashing for Cross-modal RetrievalZhan Yang, Yiran Liu, Youyuan Huang, Yinan LiAAAI 2026
- Graph Convolutional Semi-Supervised Cross-Modal HashingXiaobo Shen, Gaoyao Yu, Yinfan Chen, Xichen Yang et al.ACM MM 2024 · 5 citations
