Prompt-based Zero-shot Video Moment Retrieval
Guolong Wang, Xun Wu, Zhaoyuan Liu, Junchi Yan
Abstract
Video moment retrieval aims at localizing a specific moment from an untrimmed video by a sentence query. Most methods rely on heavy annotations of video moment-query pairs. Recent zero-shot methods reduced annotation cost, yet they neglected the global visual feature due to the separation of video and text learning process. To avoid the lack of visual features, we propose a Prompt-based Zero-shot Video Moment Retrieval (PZVMR) method. Motivated by the frame of prompt learning, we design two modules: 1) Proposal Prompt (PP): We randomly masks sequential frames to build a prompt to generate proposals; 2) Verb Prompt (VP): We provide patterns of nouns and the masked verb to build a prompt to generate pseudo queries with verbs. Our PZVMR utilizes task-relevant knowledge distilled from pre-trained CLIP and adapts the knowledge to VMR. Unlike the pioneering work, we introduce visual features into each module. Extensive experiments show that our PZVMR not only outperforms the existing zero-shot method (PSVL) on two public datasets (Charades-STA and ActivityNet-Captions) by 4.4% and 2.5% respectively in mIoU, but also outperforms several methods using stronger supervision.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8883f831-803f-4afb-b7e3-4a6778963507Cited by top-tier papers4
- Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the WildPeijun Bao, Chenqi Kong, Siyuan Yang, Zihao Shao et al.ICCV 2025 · 3 citations
- Show and Guide: Instructional-Plan Grounded Vision and Language ModelDiogo Glória-Silva, David Semedo, João MagalhãesEMNLP 2024
- HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal GroundingXinyi Xu, Hongsong Wang, Guo-Sen Xie, Caifeng Shan et al.ICLR 2026
- ResidualViT for Efficient Temporally Dense Video EncodingMattia Soldan, Fabian Caba Heilbron, Bernard Ghanem, Josef Sivic et al.ICCV 2025
Related papers
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang et al.AAAI 2020 · 170 citations
- GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment RetrievalMingyu Jeon, Sunjae Yoon, Jonghee Kim, Junyeong KimAAAI 2026 · 1 citation
- PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkMingyao Zhou, Hao Sun, Wei Xie, Ming Dong et al.NeurIPS 2025 · 1 citation
- Zero-shot Natural Language Video LocalizationJinwoo Nam, Daechul Ahn, Dongyeop Kang, Seong Jong Ha et al.ICCV 2021 · 60 citations
- Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence LocalizationMinghang Zheng, Shaogang Gong, Hailin Jin, Yuxin Peng et al.ACL 2023 · 18 citations
