Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal Retrieval
Shupeng Su, Zhisheng Zhong, Chao Zhang
Abstract
Cross-modal hashing encodes the multimedia data into a common binary hash space in which the correlations among the samples from different modalities can be effectively measured. Deep cross-modal hashing further improves the retrieval performance as the deep neural networks can generate more semantic relevant features and hash codes. In this paper, we study the unsupervised deep cross-modal hash coding and propose Deep Joint-Semantics Reconstructing Hashing (DJSRH), which has the following two main advantages. First, to learn binary codes that preserve the neighborhood structure of the original data, DJSRH constructs a novel joint-semantics affinity matrix which elaborately integrates the original neighborhood information from different modalities and accordingly is capable to capture the latent intrinsic semantic affinity for the input multi-modal instances. Second, DJSRH later trains the networks to generate binary codes that maximally reconstruct above joint-semantics relations via the proposed reconstructing framework, which is more competent for the batch-wise training as it reconstructs the specific similarity value unlike the common Laplacian constraint merely preserving the similarity order. Extensive experiments demonstrate the significant improvement by DJSRH in various cross-modal retrieval tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fe70ff8-9257-482d-89d5-06d9b7e82c9eCited by top-tier papers28
- Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal HashingJun Yu, Hao Zhou, Yibing Zhan, Dacheng TaoAAAI 2021 · 184 citations
- Cross-media Structured Common Space for Multimedia Event ExtractionManling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead et al.ACL 2020 · 87 citations
- Dual Adversarial Graph Neural Networks for Multi-label Cross-modal RetrievalShengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang et al.AAAI 2021 · 72 citations
- Progressive Spatio-Temporal Prototype Matching for Text-Video RetrievalPandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie et al.ICCV 2023 · 62 citations
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di et al.AAAI 2022 · 52 citations
Related papers
- Supervised Hierarchical Deep Hashing for Cross-Modal RetrievalYu-Wei Zhan, Xin Luo, Yongxin Wang, Xin-Shun XuACM MM 2020 · 55 citations
- Joint-modal Distribution-based Similarity Hashing for Large-scale Unsupervised Deep Cross-modal RetrievalSong Liu, Shengsheng Qian, Yang Guan, Jiawei Zhan et al.SIGIR 2020 · 214 citations
- Statistical Model-driven Similarity Hashing: Bridging Modalities for Efficient Unsupervised RetrievalMingjin Kuai, Jun Long, Zhan YangAAAI 2025 · 3 citations
- Adaptive Structural Similarity Preserving for Unsupervised Cross Modal HashingLiang Li, Baihua Zheng, Weiwei SunACM MM 2022 · 29 citations
- Graph Convolutional Multi-modal Hashing for Flexible Multimedia RetrievalXu Lu, Lei Zhu, Li Liu, Liqiang Nie et al.ACM MM 2021 · 61 citations
