Prompt-based Zero-shot Video Moment Retrieval
Guolong Wang, Xun Wu, Zhaoyuan Liu, Junchi Yan
摘要
Video moment retrieval aims at localizing a specific moment from an untrimmed video by a sentence query. Most methods rely on heavy annotations of video moment-query pairs. Recent zero-shot methods reduced annotation cost, yet they neglected the global visual feature due to the separation of video and text learning process. To avoid the lack of visual features, we propose a Prompt-based Zero-shot Video Moment Retrieval (PZVMR) method. Motivated by the frame of prompt learning, we design two modules: 1) Proposal Prompt (PP): We randomly masks sequential frames to build a prompt to generate proposals; 2) Verb Prompt (VP): We provide patterns of nouns and the masked verb to build a prompt to generate pseudo queries with verbs. Our PZVMR utilizes task-relevant knowledge distilled from pre-trained CLIP and adapts the knowledge to VMR. Unlike the pioneering work, we introduce visual features into each module. Extensive experiments show that our PZVMR not only outperforms the existing zero-shot method (PSVL) on two public datasets (Charades-STA and ActivityNet-Captions) by 4.4% and 2.5% respectively in mIoU, but also outperforms several methods using stronger supervision.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the WildPeijun Bao, Chenqi Kong, Siyuan Yang, Zihao Shao 等ICCV 2025 · 被引用 3 次
- Show and Guide: Instructional-Plan Grounded Vision and Language ModelDiogo Glória-Silva, David Semedo, João MagalhãesEMNLP 2024
- HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal GroundingXinyi Xu, Hongsong Wang, Guo-Sen Xie, Caifeng Shan 等ICLR 2026
- ResidualViT for Efficient Temporally Dense Video EncodingMattia Soldan, Fabian Caba Heilbron, Bernard Ghanem, Josef Sivic 等ICCV 2025
相关 Paper
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang 等AAAI 2020 · 被引用 170 次
- GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment RetrievalMingyu Jeon, Sunjae Yoon, Jonghee Kim, Junyeong KimAAAI 2026 · 被引用 1 次
- PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkMingyao Zhou, Hao Sun, Wei Xie, Ming Dong 等NeurIPS 2025 · 被引用 1 次
- Zero-shot Natural Language Video LocalizationJinwoo Nam, Daechul Ahn, Dongyeop Kang, Seong Jong Ha 等ICCV 2021 · 被引用 60 次
- Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence LocalizationMinghang Zheng, Shaogang Gong, Hailin Jin, Yuxin Peng 等ACL 2023 · 被引用 18 次
