3D Self-Attention for Unsupervised Video Quantization
Jingkuan Song, Ruimin Lang, Xiaosu Zhu, Xing Xu, Lianli Gao, Heng Tao Shen
摘要
Unsupervised video quantization is to compress the original videos to compact binary codes so that video retrieval can be conducted in an efficient way. In this paper, we make a first attempt to combine quantization method with video retrieval called 3D-UVQ, which obtains high retrieval accuracy with low storage cost. In the proposed framework, we address two main problems: 1) how to design an effective pipeline to perceive video contextual information for video features extraction; and 2) how to quantize these features for efficient retrieval. To tackle these problems, we propose a 3D self-attention module to exploit the spatial and temporal contextual information, where each pixel is influenced by its surrounding pixels. By taking a further recurrent operation, each pixel can finally capture the global context from all pixels. Then, we propose gradient-based residual quantization which consists of several quantization blocks to approximate the features gradually. Extensive experimental results on three benchmark datasets demonstrate that our method significantly outperforms the state-of-the-arts. Ablation study shows that both the 3D self-attention module and the gradient-based residual quantization can improve the performance of retrieval. Our model is publicly available at https://github.com/brownwolf/3D-UVQ.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Type-to-Track: Retrieve Any Object via Prompt-based TrackingPha A. Nguyen, Kha Gia Quach, Kris Kitani, Khoa LuuNeurIPS 2023 · 被引用 38 次
- A Lower Bound of Hash Codes' PerformanceXiaosu Zhu, Jingkuan Song, Yu Lei, Lianli Gao 等NeurIPS 2022 · 被引用 2 次
相关 Paper
- Neighborhood Preserving Hashing for Scalable Video RetrievalShuyan Li, Zhixiang Chen, Jiwen Lu, Xiu Li 等ICCV 2019 · 被引用 50 次
- Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure PreservationYanbin Hao, Jingru Duan, Hao Zhang, Bin Zhu 等ACM MM 2022 · 被引用 16 次
- Disentangled Representation Learning for Unsupervised Neural QuantizationHaechan Noh, Sangeek Hyun, Woojin Jeong, Hanshin Lim 等CVPR 2023
- ResQ: Residual Quantization for Video PerceptionDavide Abati, Haitam Ben Yahia, Markus Nagel, Amirhossein HabibianICCV 2023 · 被引用 3 次
- Hybrid Contrastive Quantization for Efficient Cross-View Video RetrievalJinpeng Wang, Bin Chen, Dongliang Liao, Ziyun Zeng 等WWW 2022 · 被引用 11 次
