Routing Evidence for Unseen Actions in Video Moment Retrieval
Guolong Wang, Xun Wu, Zheng Qin, Liangliang Shi
摘要
Video moment retrieval (VMR) is a cutting-edge vision-language task locating a segment in a video according to the query. Though the methods have achieved significant performance, they assume that training and testing samples share the same action types, hindering real-world application. In this paper, we specifically consider a new problem: video moment retrieval by queries with unseen actions. We propose a plug-and-play structure, Routing Evidence (RE), with multiple evidence-learning heads and dynamically route one to locate a sentence with an unseen action. Each evidence-learning head estimates the uncertainty while regressing timestamps. We formulate the evidence distribution by a Normal-Inverse Gamma function and design a router to select the most appropriate distribution for a sample. Empirically, we study the efficacy of RE on three updated databases where training and testing samples contain different action types. We find that RE outperforms other state-of-the-art methods with a more robust predictor. Code and data will be available at https://github.com/dieuroi/Routing-Evidence.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Video Temporal GroundingJian Hu, Zixu Cheng, Shaogang Gong, Isabel Guan 等NeurIPS 2025
- Optimal Flow Transport and its Entropic Regularization: a GPU-friendly Matrix Iterative Algorithm for Flow Balance SatisfactionLiangliang Shi, Yufeng Li, Kaipeng Zeng, Yihui Tu 等ICLR 2025
- SelKD: Selective Knowledge Distillation via Optimal Transport PerspectiveLiangliang Shi, Zhengyan Shi, Junchi YanICLR 2025
相关 Paper
- Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using LanguageXiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu 等ACM MM 2024 · 被引用 8 次
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou 等AAAI 2024 · 被引用 30 次
- Reverse Distribution Based Video Moment Retrieval for Effective Bias EliminationLingdu Kong, Xiaochun Yang, Tieying Li, Bin Wang 等AAAI 2025
- Semantics-Enriched Cross-Modal Alignment for Complex-Query Video Moment RetrievalXingyu Shen, Xiang Zhang, Xun Yang, Yibing Zhan 等ACM MM 2023 · 被引用 9 次
- Partial Annotation-based Video Moment Retrieval via Iterative LearningWei Ji, Renjie Liang, Lizi Liao, Hao Fei 等ACM MM 2023 · 被引用 17 次
