Lightweight Contrastive Distilled Hashing for Online Cross-modal Retrieval
Jiaxing Li, Lin Jiang, Zeqi Ma, Kaihang Jiang, Xiaozhao Fang, Jie Wen
摘要
Deep online cross-modal hashing has gained much attention from researchers recently, as its promising applications with low storage requirement, fast retrieval efficiency and cross modality adaptive, etc. However, there still exists some technical hurdles that hinder its applications, e.g., 1) how to extract the coexistent semantic relevance of cross-modal data, 2) how to achieve competitive performance when handling the real time data streams, 3) how to transfer the knowledge learned from offline to online training in a lightweight manner. To address these problems, this paper proposes a lightweight contrastive distilled hashing (LCDH) for crossmodal retrieval, by innovatively bridging the offline and online cross-modal hashing by similarity matrix approximation in a knowledge distillation framework. Specifically, in the teacher network, LCDH first extracts the cross-modal features by the contrastive language-image pre-training (CLIP), which are further fed into an attention module for representation enhancement after feature fusion. Then, the output of the attention module is fed into a FC layer to obtain hash codes for aligning the sizes of similarity matrices for online and offline training. In the student network, LCDH extracts the visual and textual features by lightweight models, and then the features are fed into a FC layer to generate binary codes. Finally, by approximating the similarity matrices, the performance of online hashing in the lightweight student network can be enhanced by the supervision of coexistent semantic relevance that is distilled from the teacher network. Experimental results on three widely used datasets demonstrate that LCDH outperforms some state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Online Collective Matrix Factorization Hashing for Large-Scale Cross-Media RetrievalDi Wang, Quan Wang, Yaqiang An, Xinbo Gao 等SIGIR 2020 · 被引用 69 次
- Label Embedding Online Hashing for Cross-Modal RetrievalYongxin Wang, Xin Luo, Xin-Shun XuACM MM 2020 · 被引用 64 次
- Dual Label-Guided Graph Refinement for Multi-View Graph ClusteringYawen Ling, Jianpeng Chen, Yazhou Ren, Xiaorong Pu 等AAAI 2023 · 被引用 53 次
- Denoising High-Order Graph ClusteringYonghao Chen, Ruibing Chen, Qiaoyun Li, Xiaozhao Fang 等ICDE 2024 · 被引用 10 次
- Creating Something From Nothing: Unsupervised Knowledge Distillation for Cross-Modal HashingHengtong Hu, Lingxi Xie, Richang Hong, Qi TianCVPR 2020
相关 Paper
- An End-To-End Graph Attention Network Hashing for Cross-Modal RetrievalHuilong Jin, Yingxue Zhang, Lei Shi, Shuang Zhang 等NeurIPS 2024 · 被引用 18 次
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 被引用 9 次
- CLIP-KD: An Empirical Study of CLIP Model DistillationChuanguang Yang, Zhulin An, Libo Huang, Junyu Bi 等CVPR 2024 · 被引用 50 次
- Unsupervised Multimodal Graph Contrastive Semantic Anchor Space Dynamic Knowledge Distillation Network for Cross-Media Hash RetrievalYang Yu, Meiyu Liang, Mengran Yin, Kangkang Lu 等ICDE 2024 · 被引用 5 次
- KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsDa Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao 等ACM MM 2025 · 被引用 10 次
