Similarity Preserving Transformer Cross-Modal Hashing for Video-Text Retrieval
Qianxin Huang, Siyao Peng, Xiaobo Shen, Yunhao Yuan, Shirui Pan
Abstract
As social networks grow exponentially, there is an increasing demand for video retrieval using natural language. Cross-modal hashing that encodes multi-modal data using compact hash code has been widely used in large-scale image-text retrieval, primarily due to its computation and storage efficiency. When applied to video-text retrieval, existing unsupervised cross-modal hashing extracts the frame- or word-level features individually, and thus ignores long-term dependencies. In addition, effectively exploiting the multi-modal structure is a remarkable challenge owing to the complex nature of video and text. To address the above issues, we propose Similarity Preserving Transformer Cross-Modal Hashing (SPTCH), a new unsupervised deep cross-modal hashing method for video-text retrieval. SPTCH encodes video and text by bidirectional transformer encoder that exploits their long-term dependencies. SPTCH constructs a multi-modal collaborative graph to model correlations among multi-modal data, and applies semantic aggregation by employing Graph Convolutional Network (GCN) on such graph. SPTCH designs unsupervised multi-modal contrastive loss and neighborhood reconstruction loss to effectively leverage inter- and intra-modal similarity structure among videos and texts. The empirical results on three video benchmark datasets illustrate that the proposed SPTCH generally outperforms state-of-the-arts in video-text retrieval.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2832cea9-f8c0-49a3-b6d3-41bf3ca7fa51Related papers
- Stationary and Clustering Transformer Hashing for Cross-modal RetrievalZhan Yang, Yiran Liu, Youyuan Huang, Yinan LiAAAI 2026
- Self-Supervised Video Hashing via Bidirectional TransformersShuyan Li, Xiu Li, Jiwen Lu, Jie ZhouCVPR 2021
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 9 citations
- Unsupervised Similarity-Fusion Transformer Hashing for Multimodal RetrievalZhan Yang, Binghong Chen, Jiajun Tang, Yinan LiACM MM 2025
- Neighborhood Preserving Hashing for Scalable Video RetrievalShuyan Li, Zhixiang Chen, Jiwen Lu, Xiu Li et al.ICCV 2019 · 50 citations
