Lune

KDD2026顶会

PROBE: VLM-Guided Discrete Structural Reconfiguration for Customized and Efficient Video Retrieval

Yiyang Gu, Kaili Liu, Tao Zhe, Binqi Chen, Jiayue Fan, Junwei Yang, Zequn Liu, Zhiping Xiao, Chong Chen, Xiao Luo, Xian-Sheng Hua, Ming Zhang

2026年份

摘要

Existing self-supervised video hashing methods achieve high efficiency by encoding videos into compact binary representations, but they typically rely on a fixed global similarity geometry that enforces a single notion of similarity across all queries. In many real-world retrieval scenarios, however, the same videos may need to be compared under different semantic criteria, such as action, scene, object, emotion, or intent. This requirement gives rise to criterion-dependent video comparison, which remains largely unexplored in efficient hashing-based retrieval frameworks. To this end, we propose an efficient VLM-guided video retrieval paradigm PROBE that leverages a frozen vision–language model (VLM) to inject rich semantic structure into a reusable hash index. Each video is encoded once into a compact binary representation, while prompts act as discrete operators that select criterion-relevant semantic subspaces within the hash space, inducing discrete structural reconfiguration of similarity geometry and enabling customized video–video comparison via efficient masked distance computation. Guided by prompt-conditioned representations from the vision–language teacher, we introduce criterion-induced subspace regularization that transfers semantic geometry into criterion-conditioned hash subspaces while preserving local relational consistency. We further employ a global unconditional alignment and reconstruction objective to stabilize the full hash space across views and datasets. A single model trained jointly on heterogeneous video datasets consistently outperforms strong video hashing baselines and generalizes effectively to unseen datasets and diverse retrieval criteria, demonstrating that VLM-guided discrete structural reconfiguration is key to efficient, customized, and reusable video retrieval. The source code is available at https://github.com/liamgu06/PROBE.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖