Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure Preservation
Yanbin Hao, Jingru Duan, Hao Zhang, Bin Zhu, Pengyuan Zhou, Xiangnan He
摘要
Unsupervised video hashing typically aims to learn a compact binary vector to represent complex video content without using manual annotations. Existing unsupervised hashing methods generally suffer from incomplete exploration of various perspective dependencies (e.g., long-range and short-range) and data structures that exist in visual contents, resulting in less discriminative hash codes. In this paper, we propose aMulti-granularity Contextualized and Multi-Structure preserved Hashing (MCMSH) method, exploring multiple axial contexts for discriminative video representation generation and various structural information for unsupervised learning simultaneously. Specifically, we delicately design three self-gating modules to separately model three granularities of dependencies (i.e., long/middle/short-range dependencies) and densely integrate them into MLP-Mixer for feature contextualization, leading to a novel model MC-MLP. To facilitate unsupervised learning, we investigate three kinds of data structures, including clusters, local neighborhood similarity structure, and inter/intra-class variations, and design a multi-objective task to train MC-MLP. These data structures show high complementarities in hash code learning. We conduct extensive experiments using three video retrieval benchmark datasets, demonstrating that our MCMSH not only boosts the performance of the backbone MLP-Mixer significantly but also outperforms the competing methods notably. Code is available at: https://github.com/haoyanbin918/MCMSH.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Contrastive Masked Autoencoders for Self-Supervised Video HashingYuting Wang, Jinpeng Wang, Bin Chen, Ziyun Zeng 等AAAI 2023 · 被引用 29 次
- CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video HashingRukai Wei, Yu Liu, Jingkuan Song, Heng Cui 等ACM MM 2023 · 被引用 15 次
- Efficient Self-Supervised Video Hashing with Selective State SpacesJinpeng Wang, Niu Lian, Jun Li, Yuting Wang 等AAAI 2025 · 被引用 7 次
- AV-NAS: Audio-Visual Multi-Level Semantic Neural Architecture Search for Video HashingYong Chen, Yuxiang Zhou, Hailiang Dong, Rui Liu 等SIGIR 2025 · 被引用 1 次
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingNiu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo 等CVPR 2025
它引用的顶会 Paper13
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 被引用 912 次
- AS-MLP: An Axial Shifted MLP Architecture for VisionDongze Lian, Zehao Yu, Xing Sun, Shenghua GaoICLR 2022 · 被引用 217 次
- Deep Unsupervised Image Hashing by Maximizing Bit EntropyYunqiang Li, Jan van GemertAAAI 2021 · 被引用 109 次
- Cross-Modality High-Frequency Transformer for MR Image Super-ResolutionChaowei Fang, Dingwen Zhang, Liang Wang, Yulun Zhang 等ACM MM 2022 · 被引用 57 次
相关 Paper
- Similarity Preserving Transformer Cross-Modal Hashing for Video-Text RetrievalQianxin Huang, Siyao Peng, Xiaobo Shen, Yunhao Yuan 等ACM MM 2024 · 被引用 1 次
- Neighborhood Preserving Hashing for Scalable Video RetrievalShuyan Li, Zhixiang Chen, Jiwen Lu, Xiu Li 等ICCV 2019 · 被引用 50 次
- Self-Supervised Video Hashing via Bidirectional TransformersShuyan Li, Xiu Li, Jiwen Lu, Jie ZhouCVPR 2021
- 3D Self-Attention for Unsupervised Video QuantizationJingkuan Song, Ruimin Lang, Xiaosu Zhu, Xing Xu 等SIGIR 2020 · 被引用 3 次
- Deep Unsupervised Hashing with Latent Semantic ComponentsQinghong Lin, Xiaojun Chen, Qin Zhang, Shaotian Cai 等AAAI 2022 · 被引用 6 次
