Video Moment Retrieval with Hierarchical Contrastive Learning
Bolin Zhang, Chao Yang, Bin Jiang, Xiaokang Zhou
Abstract
This paper explores the task of video moment retrieval (VMR), which aims to localize the temporal boundary of a specific moment from an untrimmed video by a sentence query. Previous methods either extract pre-defined candidate moment features and select the moment that best matches the query by ranking, or directly align the boundary clips of a target moment with the query and predict matching scores. Despite their effectiveness, these methods mostly focus only on aligning the query and single-level clip or moment features, and ignore the different granularities involved in the video itself, such as clip, moment, or video, resulting in insufficient cross-modal interaction. To this end, we propose a Temporal Localization Network with Hierarchical Contrastive Learning (HCLNet) for the VMR task. Specifically, we introduce a hierarchical contrastive learning method to better align the query and video by maximizing the mutual information (MI) between query and three different granularities of video to learn informative representations. Meanwhile, we introduce a self-supervised cycle-consistency loss to enforce the further semantic alignment between fine-grained video clips and query words. Experiments on three standard benchmarks show the effectiveness of our proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8e4ed108-4af2-44fc-a097-9e789d92aa48Cited by top-tier papers3
- Boosting Temporal Sentence Grounding via Causal InferenceKefan Tang, Lihuo He, Jisheng Dang, Xinbo GaoACM MM 2025 · 2 citations
- PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkMingyao Zhou, Hao Sun, Wei Xie, Ming Dong et al.NeurIPS 2025 · 1 citation
- TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight DetectionHao Sun, Mingyao Zhou, Wenjing Chen, Wei XieAAAI 2024
Related papers
- Video Corpus Moment Retrieval with Contrastive LearningHao Zhang, Aixin Sun, Wei Jing, Guoshun Nan et al.SIGIR 2021 · 88 citations
- Visual Co-Occurrence Alignment Learning for Weakly-Supervised Video Moment RetrievalZheng Wang, Jingjing Chen, Yu-Gang JiangACM MM 2021 · 74 citations
- Counterfactual Cross-modality Reasoning for Weakly Supervised Video Moment LocalizationZezhong Lv, Bing Su, Ji-Rong WenACM MM 2023 · 23 citations
- Multi-Scale Contrastive Learning for Video Temporal GroundingThong Thanh Nguyen, Yi Bin, Xiaobao Wu, Zhiyuan Hu et al.AAAI 2025 · 7 citations
- Weakly Supervised Video Moment Localization with Contrastive Negative Sample MiningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yang LiuAAAI 2022 · 109 citations
