Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using Language
Xiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Renfu Li, Zichuan Xu, Lixing Chen, Panpan Zheng, Yu Cheng
摘要
Video Moment Retrieval (VMR) targets to retrieve the specific moment corresponding to a sentence query from an untrimmed video. Although recent respectable works have made remarkable progress in this task, they implicitly are rooted in the closed-set assumption that all the given queries as video-relevant. Given an OOD query in open-set scenarios, they still utilize it for wrong retrieval, which might lead to irrecoverable losses in high-risk scenarios, e.g., criminal activity detection. To this end, we creatively explore a brand-new VMR setting termed Open-Set Video Moment Retrieval (OS-VMR), where we should not only retrieve the precise moments based on ID query, but also reject OOD queries. In this paper, we make the first attempt to step toward OS-VMR and propose a novel model OpenVMR, which first distinguishes ID and OOD queries based on the normalizing flow technology, and then conducts moment retrieval based on ID queries. Specifically, we first learn the ID distribution by constructing a normalizing flow, and assume the ID query distribution obeys the multi-variate Gaussian distribution. Then, we introduce an uncertainty score to search the ID-OOD separating boundary. After that, we refine the ID-OOD boundary by pulling together ID query features. Besides, video-query matching and frame-query matching are designed for coarse-grained and fine-grained cross-modal interaction, respectively. Finally, a positive-unlabeled learning module is introduced for moment retrieval. Experimental results on three VMR datasets show the effectiveness of our OpenVMR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Pandora's Box: Towards Building Universal Attackers against Real-World Large Vision-Language ModelsDaizong Liu, Mingyu Yang, Xiaoye Qu, Pan Zhou 等NeurIPS 2024 · 被引用 51 次
- Spotlight on Token Perception for Multimodal Reinforcement LearningSiyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo 等ICLR 2026 · 被引用 45 次
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 被引用 17 次
- Temporal Sentence Grounding with Relevance Feedback in VideosJianfeng Dong, Xiaoman Peng, Daizong Liu, Xiaoye Qu 等NeurIPS 2024 · 被引用 12 次
- Stealing Training Graphs from Graph Neural NetworksMinhua Lin, Enyan Dai, Junjie Xu, Jinyuan Jia 等KDD 2025 · 被引用 3 次
它引用的顶会 Paper71
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- Why Normalizing Flows Fail to Detect Out-of-Distribution DataPolina Kirichenko, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2020 · 被引用 370 次
- Input Complexity and Out-of-distribution Detection with Likelihood-based Generative ModelsJoan Serrà, David Álvarez, Vicenç Gómez, Olga Slizovskaia 等ICLR 2020 · 被引用 307 次
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 被引用 279 次
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
相关 Paper
- Semantics-Enriched Cross-Modal Alignment for Complex-Query Video Moment RetrievalXingyu Shen, Xiang Zhang, Xun Yang, Yibing Zhan 等ACM MM 2023 · 被引用 9 次
- Routing Evidence for Unseen Actions in Video Moment RetrievalGuolong Wang, Xun Wu, Zheng Qin, Liangliang ShiKDD 2024 · 被引用 3 次
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou 等AAAI 2024 · 被引用 30 次
- From Interference to Stability: Adversarial Reliability Correction for Video Moment Retrieval with Relevance FeedbackHao Liu, Yupeng Hu, Kun Wang, Junchao Wang 等SIGIR 2026
- Prompt-based Zero-shot Video Moment RetrievalGuolong Wang, Xun Wu, Zhaoyuan Liu, Junchi YanACM MM 2022 · 被引用 33 次
