Lightweight Contrastive Distilled Hashing for Online Cross-modal Retrieval
Jiaxing Li, Lin Jiang, Zeqi Ma, Kaihang Jiang, Xiaozhao Fang, Jie Wen
Abstract
Deep online cross-modal hashing has gained much attention from researchers recently, as its promising applications with low storage requirement, fast retrieval efficiency and cross modality adaptive, etc. However, there still exists some technical hurdles that hinder its applications, e.g., 1) how to extract the coexistent semantic relevance of cross-modal data, 2) how to achieve competitive performance when handling the real time data streams, 3) how to transfer the knowledge learned from offline to online training in a lightweight manner. To address these problems, this paper proposes a lightweight contrastive distilled hashing (LCDH) for crossmodal retrieval, by innovatively bridging the offline and online cross-modal hashing by similarity matrix approximation in a knowledge distillation framework. Specifically, in the teacher network, LCDH first extracts the cross-modal features by the contrastive language-image pre-training (CLIP), which are further fed into an attention module for representation enhancement after feature fusion. Then, the output of the attention module is fed into a FC layer to obtain hash codes for aligning the sizes of similarity matrices for online and offline training. In the student network, LCDH extracts the visual and textual features by lightweight models, and then the features are fed into a FC layer to generate binary codes. Finally, by approximating the similarity matrices, the performance of online hashing in the lightweight student network can be enhanced by the supervision of coexistent semantic relevance that is distilled from the teacher network. Experimental results on three widely used datasets demonstrate that LCDH outperforms some state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e06adf15-0e97-4278-a817-ba02720f5cb6Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Online Collective Matrix Factorization Hashing for Large-Scale Cross-Media RetrievalDi Wang, Quan Wang, Yaqiang An, Xinbo Gao et al.SIGIR 2020 · 69 citations
- Label Embedding Online Hashing for Cross-Modal RetrievalYongxin Wang, Xin Luo, Xin-Shun XuACM MM 2020 · 64 citations
- Dual Label-Guided Graph Refinement for Multi-View Graph ClusteringYawen Ling, Jianpeng Chen, Yazhou Ren, Xiaorong Pu et al.AAAI 2023 · 53 citations
- Denoising High-Order Graph ClusteringYonghao Chen, Ruibing Chen, Qiaoyun Li, Xiaozhao Fang et al.ICDE 2024 · 10 citations
- Creating Something From Nothing: Unsupervised Knowledge Distillation for Cross-Modal HashingHengtong Hu, Lingxi Xie, Richang Hong, Qi TianCVPR 2020
Related papers
- An End-To-End Graph Attention Network Hashing for Cross-Modal RetrievalHuilong Jin, Yingxue Zhang, Lei Shi, Shuang Zhang et al.NeurIPS 2024 · 18 citations
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 9 citations
- CLIP-KD: An Empirical Study of CLIP Model DistillationChuanguang Yang, Zhulin An, Libo Huang, Junyu Bi et al.CVPR 2024 · 50 citations
- Unsupervised Multimodal Graph Contrastive Semantic Anchor Space Dynamic Knowledge Distillation Network for Cross-Media Hash RetrievalYang Yu, Meiyu Liang, Mengran Yin, Kangkang Lu et al.ICDE 2024 · 5 citations
- KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsDa Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao et al.ACM MM 2025 · 10 citations
