Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding
Zhenzhi Wang, Limin Wang, Tao Wu, Tianhao Li, Gangshan Wu
Abstract
Temporal grounding aims to localize a video moment which is semantically aligned with a given natural language query. Existing methods typically apply a detection or regression pipeline on the fused representation with the research focus on designing complicated prediction heads or fusion strategies. Instead, from a perspective on temporal grounding as a metric-learning problem, we present a Mutual Matching Network (MMN), to directly model the similarity between language queries and video moments in a joint embedding space. This new metric-learning framework enables fully exploiting negative samples from two new aspects: constructing negative cross-modal pairs in a mutual matching scheme and mining negative pairs across different videos. These new negative samples could enhance the joint representation learning of two modalities via cross-modal mutual matching to maximize their mutual information. Experiments show that our MMN achieves highly competitive performance compared with the state-of-the-art methods on four video grounding benchmarks. Based on MMN, we present a winner solution for the HC-STVG challenge of the 3rd PIC workshop. This suggests that metric learning is still a promising method for temporal grounding via capturing the essential cross-modal correlation in a joint embedding space. Code is available at https://github.com/MCG-NJU/MMN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81641b37-3b01-4565-911c-5fc30aee8cfeCited by top-tier papers37
- MomentDiff: Generative Video Moment Retrieval from Random to RealPandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao et al.NeurIPS 2023 · 113 citations
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment RetrievalZhihang Liu, Jun Li, Hongtao Xie, Pandeng Li et al.AAAI 2024 · 49 citations
- G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game TheoryHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li et al.ICCV 2023 · 32 citations
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou et al.AAAI 2024 · 30 citations
Builds on18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
Related papers
- Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video GroundingZihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin et al.CVPR 2023
- Dual Path Interaction Network for Video Moment LocalizationHao Wang, Zheng-Jun Zha, Xuejin Chen, Zhiwei Xiong et al.ACM MM 2020 · 69 citations
- Moment Quantization for Video Temporal GroundingXiaolong Sun, Le Wang, Sanping Zhou, Liushuai Shi et al.ICCV 2025 · 2 citations
- Efficient Spatio-Temporal Video Grounding with Semantic-Guided Feature DecompositionWeikang Wang, Jing Liu, Yuting Su, Weizhi NieACM MM 2023 · 8 citations
- CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal GroundingZhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao et al.ACL 2023 · 17 citations
