Differentiable Cross-modal Hashing via Multimodal Transformers
Junfeng Tu, Xueliang Liu, Zongxiang Lin, Richang Hong, Meng Wang
摘要
Cross-modal hashing aims at projecting the cross modal content into a common Hamming space for efficient search. Most existing work first encodes the samples with a deep network and then binaries the encoded feature into hashing code. However, the relative location information in the image may be lost when an image is encoded by the convolutional network, which makes it challenging to model the relationship of different modalities. Moreover, it is NP-hard to optimize the model with the discrete sign binary function popularly used in existing solutions. To address these issues, we propose a differentiable cross-modal hashing method that utilizes the multimodal transformer as the backbone to capture the location information in an image when encoding the visual content. In addition, a novel differentiable cross-modal hashing method is proposed to generate the binary code by a selecting mechanism, which could be formulated as a continuous and easily optimized problem. We perform extensive experiments on several cross modal datasets and the results show that the proposed method outperforms many existing solutions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper14
- An End-To-End Graph Attention Network Hashing for Cross-Modal RetrievalHuilong Jin, Yingxue Zhang, Lei Shi, Shuang Zhang 等NeurIPS 2024 · 被引用 18 次
- Pairwise-Label-Based Deep Incremental Hashing with Simultaneous Code ExpansionDayan Wu, Qinghang Su, Bo Li, Weiping WangAAAI 2024 · 被引用 11 次
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 被引用 9 次
- Exploiting Descriptive Completeness Prior for Cross Modal Hashing with Incomplete LabelsHaoyang Luo, Zheng Zhang, Yadan LuoNeurIPS 2024 · 被引用 5 次
相关 Paper
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 被引用 261 次
- Multi-Granularity Interactive Transformer Hashing for Cross-modal RetrievalYishu Liu, Qingpeng Wu, Zheng Zhang, Jingyi Zhang 等ACM MM 2023 · 被引用 44 次
- Supervised Hierarchical Deep Hashing for Cross-Modal RetrievalYu-Wei Zhan, Xin Luo, Yongxin Wang, Xin-Shun XuACM MM 2020 · 被引用 55 次
- Stationary and Clustering Transformer Hashing for Cross-modal RetrievalZhan Yang, Yiran Liu, Youyuan Huang, Yinan LiAAAI 2026
- Graph Convolutional Semi-Supervised Cross-Modal HashingXiaobo Shen, Gaoyao Yu, Yinfan Chen, Xichen Yang 等ACM MM 2024 · 被引用 5 次
