Million-Scale Text-to-Video Retrieval with Hyperdimensional Computing
Hyunsei Lee, Jaewoo Gwak, Shinhyoung Jang, Junyoung Lee, Yeseong Kim
Abstract
Scalable video retrieval is increasingly challenging as datasets reach tens of millions of videos. Current text-to-video retrieval (T2VR) methods either compress videos into single dense vectors, losing segment-level detail, or expand them into multi-frame representations, incurring prohibitive storage and search costs. We propose a binary hyperdimensional representation that encodes each video into a compact 3,072-dimension hypervector, preserving semantic fidelity while reducing memory via bit-packing. To leverage the properties of hypervectors for sublinear search, we introduce Hypervector Retrieval (HVR), a frequency-aware inverted index that prioritizes rare informative positions and refines candidates using GPU-accelerated Hamming search. Experiments show that our approach matches or exceeds dense baselines for T2VR and surpasses state-of-the-art partially relevant video retrieval (PRVR) by over 5% Recall@ 10 on ActivityNet. At scale, HVR processes over 2,000 queries per second on 10M videos, maintains recall within 1% of exact search, and achieves 5.3× greater storage capacity than CLIP4Clip and over 2,116× over MS-SL.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang et al.ACM MM 2022 · 65 citations
- Prototypes Are Balanced Units for Efficient and Effective Partially Relevant Video RetrievalWonJun Moon, Cheol-Ho Cho, Woojin Jun, Taeoh Kim et al.ICCV 2025 · 3 citations
- Holistic Features are Almost Sufficient for Text-to-Video RetrievalKaibin Tian, Ruixiang Zhao, Zijie Xin, Bangxiang Lan et al.CVPR 2024 · 15 citations
- GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video RetrievalYuting Wang, Jinpeng Wang, Bin Chen, Ziyun Zeng et al.AAAI 2024 · 32 citations
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late InteractionZilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen et al.ICLR 2026 · 40 citations
