Self-Supervised Video Hashing via Bidirectional Transformers
Shuyan Li, Xiu Li, Jiwen Lu, Jie Zhou
Abstract
Most existing unsupervised video hashing methods are built on unidirectional models with less reliable training objectives, which underuse the correlations among frames and the similarity structure between videos. To enable efficient scalable video retrieval, we propose a self-supervised video Hashing method based on Bidirectional Transformers (BTH). Based on the encoder-decoder structure of transformers, we design a visual cloze task to fully exploit the bidirectional correlations between frames. To unveil the similarity structure between unlabeled video data, we further develop a similarity reconstruction task by establishing reliable and effective similarity connections in the video space. Furthermore, we develop a cluster assignment task to exploit the structural statistics of the whole dataset such that more discriminative binary codes can be learned. Extensive experiments implemented on three public benchmark datasets, FCVID, ActivityNet and YFCC, demonstrate the superiority of our proposed approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4cf84cf-68c8-42ec-90e3-77107f134841Cited by top-tier papers6
- Contrastive Masked Autoencoders for Self-Supervised Video HashingYuting Wang, Jinpeng Wang, Bin Chen, Ziyun Zeng et al.AAAI 2023 · 29 citations
- Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure PreservationYanbin Hao, Jingru Duan, Hao Zhang, Bin Zhu et al.ACM MM 2022 · 16 citations
- CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video HashingRukai Wei, Yu Liu, Jingkuan Song, Heng Cui et al.ACM MM 2023 · 15 citations
- Efficient Self-Supervised Video Hashing with Selective State SpacesJinpeng Wang, Niu Lian, Jun Li, Yuting Wang et al.AAAI 2025 · 7 citations
- AV-NAS: Audio-Visual Multi-Level Semantic Neural Architecture Search for Video HashingYong Chen, Yuxiang Zhou, Hailiang Dong, Rui Liu et al.SIGIR 2025 · 1 citation
Builds on2
Related papers
- Similarity Preserving Transformer Cross-Modal Hashing for Video-Text RetrievalQianxin Huang, Siyao Peng, Xiaobo Shen, Yunhao Yuan et al.ACM MM 2024 · 1 citation
- AVHash: Joint Audio-Visual Hashing for Video RetrievalYuxiang Zhou, Zhe Sun, Rui Liu, Yong Chen et al.ACM MM 2024 · 4 citations
- Unsupervised Similarity-Fusion Transformer Hashing for Multimodal RetrievalZhan Yang, Binghong Chen, Jiajun Tang, Yinan LiACM MM 2025
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingNiu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo et al.CVPR 2025
- Stationary and Clustering Transformer Hashing for Cross-modal RetrievalZhan Yang, Yiran Liu, Youyuan Huang, Yinan LiAAAI 2026
