Visual Co-Occurrence Alignment Learning for Weakly-Supervised Video Moment Retrieval
Zheng Wang, Jingjing Chen, Yu-Gang Jiang
Abstract
Video moment retrieval aims to localize the most relevant video moment given the text query. Weakly supervised approaches leverage video-text pairs only for training, without temporal annotations. Most current methods align the proposed video moment and the text in a joint embedding space. However, in lack of temporal annotations, the semantic gap between these two modalities makes it predominant to learn joint feature representation for most methods, with less emphasis on learning visual feature representation. This paper aims to improve the visual feature representation with supervisions in the visual domain, obtaining discriminative visual features for cross-modal learning. Based on the observation that relevant video moments (i.e., share similar activities) from different videos are commonly described by similar sentences; hence the visual features of these relevant video moments should also be similar despite that they come from different videos. Therefore, to obtain more discriminative and robust visual features for video moment retrieval, we propose to align the visual features of relevant video moments from different videos that co-occurred in the same training batch. Besides, a contrastive learning approach is introduced for learning the moment-level alignment of these videos. Through extensive experiments, we demonstrate that the proposed visual co-occurrence alignment learning method outperforms the cross-modal alignment learning counterpart and achieves promising results for video moment retrieval.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e0a613a9-35d6-46dc-a76b-648eb212628dCited by top-tier papers22
- Balanced Contrastive Learning for Long-Tailed Visual RecognitionJianggang Zhu, Zheng Wang, Jingjing Chen, Yi-Ping Phoebe Chen et al.CVPR 2022 · 194 citations
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng et al.CVPR 2022 · 108 citations
- Hypotheses Tree Building for One-Shot Temporal Sentence LocalizationDaizong Liu, Xiang Fang, Pan Zhou, Xing Di et al.AAAI 2023 · 29 citations
- Attacking Video Recognition Models with Bullet-Screen CommentsKai Chen, Zhipeng Wei, Jingjing Chen, Zuxuan Wu et al.AAAI 2022 · 27 citations
- Curriculum Multi-Negative Augmentation for Debiased Video GroundingXiaohan Lan, Yitian Yuan, Hong Chen, Xin Wang et al.AAAI 2023 · 26 citations
Related papers
- Counterfactual Cross-modality Reasoning for Weakly Supervised Video Moment LocalizationZezhong Lv, Bing Su, Ji-Rong WenACM MM 2023 · 23 citations
- Multi-Modal Relational Graph for Cross-Modal Video Moment RetrievalYawen Zeng, Da Cao, Xiaochi Wei, Meng Liu et al.CVPR 2021
- Video Corpus Moment Retrieval with Contrastive LearningHao Zhang, Aixin Sun, Wei Jing, Guoshun Nan et al.SIGIR 2021 · 88 citations
- Cross-Sentence Temporal and Semantic Relations in Video Activity LocalisationJiabo Huang, Yang Liu, Shaogang Gong, Hailin JinICCV 2021 · 77 citations
- Video Moment Retrieval with Hierarchical Contrastive LearningBolin Zhang, Chao Yang, Bin Jiang, Xiaokang ZhouACM MM 2022 · 21 citations
