Lune

EuroSys2026Top-tier venue

Million-Scale Text-to-Video Retrieval with Hyperdimensional Computing

Hyunsei Lee, Jaewoo Gwak, Shinhyoung Jang, Junyoung Lee, Yeseong Kim

2026Year

Abstract

Scalable video retrieval is increasingly challenging as datasets reach tens of millions of videos. Current text-to-video retrieval (T2VR) methods either compress videos into single dense vectors, losing segment-level detail, or expand them into multi-frame representations, incurring prohibitive storage and search costs. We propose a binary hyperdimensional representation that encodes each video into a compact 3,072-dimension hypervector, preserving semantic fidelity while reducing memory via bit-packing. To leverage the properties of hypervectors for sublinear search, we introduce Hypervector Retrieval (HVR), a frequency-aware inverted index that prioritizes rare informative positions and refines candidates using GPU-accelerated Hamming search. Experiments show that our approach matches or exceeds dense baselines for T2VR and surpasses state-of-the-art partially relevant video retrieval (PRVR) by over 5% Recall@ 10 on ActivityNet. At scale, HVR processes over 2,000 queries per second on 10M videos, maintains recall within 1% of exact search, and achieves 5.3× greater storage capacity than CLIP4Clip and over 2,116× over MS-SL.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines