PROBE: VLM-Guided Discrete Structural Reconfiguration for Customized and Efficient Video Retrieval
Yiyang Gu, Kaili Liu, Tao Zhe, Binqi Chen, Jiayue Fan, Junwei Yang, Zequn Liu, Zhiping Xiao, Chong Chen, Xiao Luo, Xian-Sheng Hua, Ming Zhang
Abstract
Existing self-supervised video hashing methods achieve high efficiency by encoding videos into compact binary representations, but they typically rely on a fixed global similarity geometry that enforces a single notion of similarity across all queries. In many real-world retrieval scenarios, however, the same videos may need to be compared under different semantic criteria, such as action, scene, object, emotion, or intent. This requirement gives rise to criterion-dependent video comparison, which remains largely unexplored in efficient hashing-based retrieval frameworks. To this end, we propose an efficient VLM-guided video retrieval paradigm PROBE that leverages a frozen vision–language model (VLM) to inject rich semantic structure into a reusable hash index. Each video is encoded once into a compact binary representation, while prompts act as discrete operators that select criterion-relevant semantic subspaces within the hash space, inducing discrete structural reconfiguration of similarity geometry and enabling customized video–video comparison via efficient masked distance computation. Guided by prompt-conditioned representations from the vision–language teacher, we introduce criterion-induced subspace regularization that transfers semantic geometry into criterion-conditioned hash subspaces while preserving local relational consistency. We further employ a global unconditional alignment and reconstruction objective to stabilize the full hash space across views and datasets. A single model trained jointly on heterogeneous video datasets consistently outperforms strong video hashing baselines and generalizes effectively to unseen datasets and diverse retrieval criteria, demonstrating that VLM-guided discrete structural reconfiguration is key to efficient, customized, and reusable video retrieval. The source code is available at https://github.com/liamgu06/PROBE.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 51f93ce5-9981-436f-88f7-ab3983f244eaRelated papers
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingNiu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo et al.CVPR 2025
- Contrastive Masked Autoencoders for Self-Supervised Video HashingYuting Wang, Jinpeng Wang, Bin Chen, Ziyun Zeng et al.AAAI 2023 · 29 citations
- Self-Supervised Video Hashing via Bidirectional TransformersShuyan Li, Xiu Li, Jiwen Lu, Jie ZhouCVPR 2021
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 9 citations
- VTD-CLIP: Video-to-Text Discretization via Prompting CLIPWencheng Zhu, Yuexin Wang, Hongxuan Li, Pengfei ZhuAAAI 2026 · 2 citations
