Video Moment Retrieval from Text Queries via Single Frame Annotation
Ran Cui, Tianwen Qian, Pai Peng, Elena Daskalaki, Jingjing Chen, Xiaowei Guo, Huyang Sun, Yu-Gang Jiang
摘要
Video moment retrieval aims at finding the start and end timestamps of a moment (part of a video) described by a given natural language query. Fully supervised methods need complete temporal boundary annotations to achieve promising results, which is costly since the annotator needs to watch the whole moment. Weakly supervised methods only rely on the paired video and query, but the performance is relatively poor. In this paper, we look closer into the annotation process and propose a new paradigm called "glance annotation". This paradigm requires the timestamp of only one single random frame, which we refer to as a "glance", within the temporal boundary of the fully supervised counterpart. We argue this is beneficial because comparing to weak supervision, trivial cost is added yet more potential in performance is provided. Under the glance annotation setting, we propose a method named as Video moment retrieval via Glance Annotation (ViGA) 1 based on contrastive learning. ViGA cuts the input video into clips and contrasts between clips and queries, in which glance guided Gaussian distributed weights are assigned to all clips. Our extensive experiments indicate that ViGA achieves better results than the state-of-the-art weakly supervised methods by a large margin, even comparable to fully supervised methods in some cases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 被引用 281 次
- Hypotheses Tree Building for One-Shot Temporal Sentence LocalizationDaizong Liu, Xiang Fang, Pan Zhou, Xing Di 等AAAI 2023 · 被引用 29 次
- Faster Video Moment Retrieval with Point-Level SupervisionXun Jiang, Zailei Zhou, Xing Xu, Yang Yang 等ACM MM 2023 · 被引用 24 次
- D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance AnnotationHanjun Li, Xiujun Shu, Sunan He, Ruizhi Qiao 等ICCV 2023 · 被引用 21 次
- Partial Annotation-based Video Moment Retrieval via Iterative LearningWei Ji, Renjie Liang, Lizi Liao, Hao Fei 等ACM MM 2023 · 被引用 17 次
它引用的顶会 Paper11
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 被引用 279 次
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 被引用 206 次
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang 等AAAI 2020 · 被引用 170 次
- Fine-grained Iterative Attention Network for Temporal Language Localization in VideosXiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng 等ACM MM 2020 · 被引用 92 次
相关 Paper
- Weakly Supervised Video Moment Localization with Contrastive Negative Sample MiningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yang LiuAAAI 2022 · 被引用 109 次
- PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkMingyao Zhou, Hao Sun, Wei Xie, Ming Dong 等NeurIPS 2025 · 被引用 1 次
- Visual Co-Occurrence Alignment Learning for Weakly-Supervised Video Moment RetrievalZheng Wang, Jingjing Chen, Yu-Gang JiangACM MM 2021 · 被引用 74 次
- Counterfactual Cross-modality Reasoning for Weakly Supervised Video Moment LocalizationZezhong Lv, Bing Su, Ji-Rong WenACM MM 2023 · 被引用 23 次
- Video Moment Retrieval with Hierarchical Contrastive LearningBolin Zhang, Chao Yang, Bin Jiang, Xiaokang ZhouACM MM 2022 · 被引用 21 次
