Contrastive Masked Autoencoders for Self-Supervised Video Hashing
Yuting Wang, Jinpeng Wang, Bin Chen, Ziyun Zeng, Shu-Tao Xia
Abstract
Self-Supervised Video Hashing (SSVH) models learn to generate short binary representations for videos without ground-truth supervision, facilitating large-scale video retrieval efficiency and attracting increasing research attention. The success of SSVH lies in the understanding of video content and the ability to capture the semantic relation among unlabeled videos. Typically, state-of-the-art SSVH methods consider these two points in a two-stage training pipeline, where they firstly train an auxiliary network by instance-wise mask-and-predict tasks and secondly train a hashing model to preserve the pseudo-neighborhood structure transferred from the auxiliary network. This consecutive training strategy is inflexible and also unnecessary. In this paper, we propose a simple yet effective one-stage SSVH method called ConMH, which incorporates video semantic information and video similarity relationship understanding in a single stage. To capture video semantic information for better hashing learning, we adopt an encoder-decoder structure to reconstruct the video from its temporal-masked frames. Particularly, we find that a higher masking ratio helps video understanding. Besides, we fully exploit the similarity relationship between videos by maximizing agreement between two augmented views of a video, which contributes to more discriminative and robust hash codes. Extensive experiments on three large-scale video datasets (i.e., FCVID, ActivityNet and YFCC) indicate that ConMH achieves state-of-the-art results. Code is available at https://github.com/huangmozhi9527/ConMH.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 663c3e47-98e8-4604-bc5b-3e6cbb366144Cited by top-tier papers8
- GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video RetrievalYuting Wang, Jinpeng Wang, Bin Chen, Ziyun Zeng et al.AAAI 2024 · 32 citations
- CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video HashingRukai Wei, Yu Liu, Jingkuan Song, Heng Cui et al.ACM MM 2023 · 15 citations
- Efficient Self-Supervised Video Hashing with Selective State SpacesJinpeng Wang, Niu Lian, Jun Li, Yuting Wang et al.AAAI 2025 · 7 citations
- Generalized Debiased Semi-Supervised Hashing for Large-Scale Image RetrievalXingbo Liu, Xuening Zhang, Xiushan Nie, Yang Shi et al.AAAI 2025 · 4 citations
- InfoMAE: Pair-Efficient Cross-Modal Alignment for Multimodal Time-Series Sensing SignalsTomoyoshi Kimura, Xinlin Li, Osama A. Hanna, Yatong Chen et al.WWW 2025 · 3 citations
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingNiu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo et al.CVPR 2025
- Self-Supervised Video Hashing via Bidirectional TransformersShuyan Li, Xiu Li, Jiwen Lu, Jie ZhouCVPR 2021
- Neighborhood Preserving Hashing for Scalable Video RetrievalShuyan Li, Zhixiang Chen, Jiwen Lu, Xiu Li et al.ICCV 2019 · 50 citations
- Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure PreservationYanbin Hao, Jingru Duan, Hao Zhang, Bin Zhu et al.ACM MM 2022 · 16 citations
- Exposing the Self-Supervised Space-Time Correspondence Learning via Graph KernelsZheyun Qin, Xiankai Lu, Xiushan Nie, Yilong Yin et al.AAAI 2023 · 22 citations
