Hybrid Contrastive Quantization for Efficient Cross-View Video Retrieval
Jinpeng Wang, Bin Chen, Dongliang Liao, Ziyun Zeng, Gongfu Li, Shu-Tao Xia, Jin Xu
Abstract
With the recent boom of video-based social platforms (e.g., YouTube and TikTok), video retrieval using sentence queries has become an important demand and attracts increasing research attention. Despite the decent performance, existing text-video retrieval models in vision and language communities are impractical for large-scale Web search because they adopt brute-force search based on high-dimensional embeddings. To improve efficiency, Web search engines widely apply vector compression libraries (e.g., FAISS [26]) to post-process the learned embeddings. Unfortunately, separate compression from feature encoding degrades the robustness of representations and incurs performance decay. To pursue a better balance between performance and efficiency, we propose the first quantized representation learning method for cross-view video retrieval, namely Hybrid Contrastive Quantization (HCQ). Specifically, HCQ learns both coarse-grained and fine-grained quantizations with transformers, which provide complementary understandings for texts and videos and preserve comprehensive semantic information. By performing Asymmetric-Quantized Contrastive Learning (AQ-CL) across views, HCQ aligns texts and videos at coarse-grained and multiple fine-grained levels. This hybrid-grained learning strategy serves as strong supervision on the cross-view video quantization model, where contrastive learning at different levels can be mutually promoted. Extensive experiments on three Web video benchmark datasets demonstrate that HCQ achieves competitive performance with state-of-the-art non-compressed retrieval methods while showing high efficiency in storage and computation. Code and configurations are available at https://github.com/gimpong/WWW22-HCQ.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0b2f19f-03e5-4564-a2d2-a5219ce881b1Cited by top-tier papers4
- Enhancing Partially Relevant Video Retrieval with Hyperbolic LearningJun Li, Jinpeng Wang, Chaolei Tan, Niu Lian et al.ICCV 2025 · 5 citations
- EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language ModelsGuanghao Meng, Sunan He, Jinpeng Wang, Tao Dai et al.AAAI 2025 · 5 citations
- Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic LearningHaomiao Tang, Jinpeng Wang, Minyi Zhao, Guanghao Meng et al.AAAI 2026 · 1 citation
- Imagine with Layout and Sketch: Enhancing Vision-Language Retrieval with Dual-Stream Multi-Modal Query RefinementGuanghao Meng, Jinpeng Wang, Qian-Wei Wang, Xudong Ren et al.AAAI 2026 · 1 citation
Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningSimon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxNeurIPS 2020 · 186 citations
- Fine-Grained Action Retrieval Through Multiple Parts-of-Speech EmbeddingsMichael Wray, Gabriela Csurka, Diane Larlus, Dima DamenICCV 2019 · 185 citations
- HiT: Hierarchical Transformer with Momentum Contrast for Video-Text RetrievalSong Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen et al.ICCV 2021 · 172 citations
Related papers
- Towards Fast Adaptation of Pretrained Contrastive Models for Multi-channel Video-Language RetrievalXudong Lin, Simran Tiwari, Shiyuan Huang, Manling Li et al.CVPR 2023
- Video Corpus Moment Retrieval with Contrastive LearningHao Zhang, Aixin Sun, Wei Jing, Guoshun Nan et al.SIGIR 2021 · 88 citations
- Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation LearningKaibin Tian, Yanhua Cheng, Yi Liu, Xinglin Hou et al.AAAI 2024 · 19 citations
- Similarity Preserving Transformer Cross-Modal Hashing for Video-Text RetrievalQianxin Huang, Siyao Peng, Xiaobo Shen, Yunhao Yuan et al.ACM MM 2024 · 1 citation
- 3D Self-Attention for Unsupervised Video QuantizationJingkuan Song, Ruimin Lang, Xiaosu Zhu, Xing Xu et al.SIGIR 2020 · 3 citations
