Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure Preservation
Yanbin Hao, Jingru Duan, Hao Zhang, Bin Zhu, Pengyuan Zhou, Xiangnan He
Abstract
Unsupervised video hashing typically aims to learn a compact binary vector to represent complex video content without using manual annotations. Existing unsupervised hashing methods generally suffer from incomplete exploration of various perspective dependencies (e.g., long-range and short-range) and data structures that exist in visual contents, resulting in less discriminative hash codes. In this paper, we propose aMulti-granularity Contextualized and Multi-Structure preserved Hashing (MCMSH) method, exploring multiple axial contexts for discriminative video representation generation and various structural information for unsupervised learning simultaneously. Specifically, we delicately design three self-gating modules to separately model three granularities of dependencies (i.e., long/middle/short-range dependencies) and densely integrate them into MLP-Mixer for feature contextualization, leading to a novel model MC-MLP. To facilitate unsupervised learning, we investigate three kinds of data structures, including clusters, local neighborhood similarity structure, and inter/intra-class variations, and design a multi-objective task to train MC-MLP. These data structures show high complementarities in hash code learning. We conduct extensive experiments using three video retrieval benchmark datasets, demonstrating that our MCMSH not only boosts the performance of the backbone MLP-Mixer significantly but also outperforms the competing methods notably. Code is available at: https://github.com/haoyanbin918/MCMSH.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Contrastive Masked Autoencoders for Self-Supervised Video HashingYuting Wang, Jinpeng Wang, Bin Chen, Ziyun Zeng et al.AAAI 2023 · 29 citations
- CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video HashingRukai Wei, Yu Liu, Jingkuan Song, Heng Cui et al.ACM MM 2023 · 15 citations
- Efficient Self-Supervised Video Hashing with Selective State SpacesJinpeng Wang, Niu Lian, Jun Li, Yuting Wang et al.AAAI 2025 · 7 citations
- AV-NAS: Audio-Visual Multi-Level Semantic Neural Architecture Search for Video HashingYong Chen, Yuxiang Zhou, Hailiang Dong, Rui Liu et al.SIGIR 2025 · 1 citation
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingNiu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo et al.CVPR 2025
Builds on13
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 912 citations
- AS-MLP: An Axial Shifted MLP Architecture for VisionDongze Lian, Zehao Yu, Xing Sun, Shenghua GaoICLR 2022 · 217 citations
- Deep Unsupervised Image Hashing by Maximizing Bit EntropyYunqiang Li, Jan van GemertAAAI 2021 · 109 citations
- Cross-Modality High-Frequency Transformer for MR Image Super-ResolutionChaowei Fang, Dingwen Zhang, Liang Wang, Yulun Zhang et al.ACM MM 2022 · 57 citations
Related papers
- Similarity Preserving Transformer Cross-Modal Hashing for Video-Text RetrievalQianxin Huang, Siyao Peng, Xiaobo Shen, Yunhao Yuan et al.ACM MM 2024 · 1 citation
- Neighborhood Preserving Hashing for Scalable Video RetrievalShuyan Li, Zhixiang Chen, Jiwen Lu, Xiu Li et al.ICCV 2019 · 50 citations
- Self-Supervised Video Hashing via Bidirectional TransformersShuyan Li, Xiu Li, Jiwen Lu, Jie ZhouCVPR 2021
- 3D Self-Attention for Unsupervised Video QuantizationJingkuan Song, Ruimin Lang, Xiaosu Zhu, Xing Xu et al.SIGIR 2020 · 3 citations
- Deep Unsupervised Hashing with Latent Semantic ComponentsQinghong Lin, Xiaojun Chen, Qin Zhang, Shaotian Cai et al.AAAI 2022 · 6 citations
