PROBE: VLM-Guided Discrete Structural Reconfiguration for Customized and Efficient Video Retrieval
Yiyang Gu, Kaili Liu, Tao Zhe, Binqi Chen, Jiayue Fan, Junwei Yang, Zequn Liu, Zhiping Xiao, Chong Chen, Xiao Luo, Xian-Sheng Hua, Ming Zhang
摘要
Existing self-supervised video hashing methods achieve high efficiency by encoding videos into compact binary representations, but they typically rely on a fixed global similarity geometry that enforces a single notion of similarity across all queries. In many real-world retrieval scenarios, however, the same videos may need to be compared under different semantic criteria, such as action, scene, object, emotion, or intent. This requirement gives rise to criterion-dependent video comparison, which remains largely unexplored in efficient hashing-based retrieval frameworks. To this end, we propose an efficient VLM-guided video retrieval paradigm PROBE that leverages a frozen vision–language model (VLM) to inject rich semantic structure into a reusable hash index. Each video is encoded once into a compact binary representation, while prompts act as discrete operators that select criterion-relevant semantic subspaces within the hash space, inducing discrete structural reconfiguration of similarity geometry and enabling customized video–video comparison via efficient masked distance computation. Guided by prompt-conditioned representations from the vision–language teacher, we introduce criterion-induced subspace regularization that transfers semantic geometry into criterion-conditioned hash subspaces while preserving local relational consistency. We further employ a global unconditional alignment and reconstruction objective to stabilize the full hash space across views and datasets. A single model trained jointly on heterogeneous video datasets consistently outperforms strong video hashing baselines and generalizes effectively to unseen datasets and diverse retrieval criteria, demonstrating that VLM-guided discrete structural reconfiguration is key to efficient, customized, and reusable video retrieval. The source code is available at https://github.com/liamgu06/PROBE.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingNiu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo 等CVPR 2025
- Contrastive Masked Autoencoders for Self-Supervised Video HashingYuting Wang, Jinpeng Wang, Bin Chen, Ziyun Zeng 等AAAI 2023 · 被引用 29 次
- Self-Supervised Video Hashing via Bidirectional TransformersShuyan Li, Xiu Li, Jiwen Lu, Jie ZhouCVPR 2021
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 被引用 9 次
- VTD-CLIP: Video-to-Text Discretization via Prompting CLIPWencheng Zhu, Yuexin Wang, Hongxuan Li, Pengfei ZhuAAAI 2026 · 被引用 2 次
