Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding
Zhenzhi Wang, Limin Wang, Tao Wu, Tianhao Li, Gangshan Wu
摘要
Temporal grounding aims to localize a video moment which is semantically aligned with a given natural language query. Existing methods typically apply a detection or regression pipeline on the fused representation with the research focus on designing complicated prediction heads or fusion strategies. Instead, from a perspective on temporal grounding as a metric-learning problem, we present a Mutual Matching Network (MMN), to directly model the similarity between language queries and video moments in a joint embedding space. This new metric-learning framework enables fully exploiting negative samples from two new aspects: constructing negative cross-modal pairs in a mutual matching scheme and mining negative pairs across different videos. These new negative samples could enhance the joint representation learning of two modalities via cross-modal mutual matching to maximize their mutual information. Experiments show that our MMN achieves highly competitive performance compared with the state-of-the-art methods on four video grounding benchmarks. Based on MMN, we present a winner solution for the HC-STVG challenge of the 3rd PIC workshop. This suggests that metric learning is still a promising method for temporal grounding via capturing the essential cross-modal correlation in a joint embedding space. Code is available at https://github.com/MCG-NJU/MMN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- MomentDiff: Generative Video Moment Retrieval from Random to RealPandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao 等NeurIPS 2023 · 被引用 113 次
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun 等CVPR 2024 · 被引用 83 次
- Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment RetrievalZhihang Liu, Jun Li, Hongtao Xie, Pandeng Li 等AAAI 2024 · 被引用 49 次
- G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game TheoryHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li 等ICCV 2023 · 被引用 32 次
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou 等AAAI 2024 · 被引用 30 次
它引用的顶会 Paper18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
相关 Paper
- Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video GroundingZihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin 等CVPR 2023
- Dual Path Interaction Network for Video Moment LocalizationHao Wang, Zheng-Jun Zha, Xuejin Chen, Zhiwei Xiong 等ACM MM 2020 · 被引用 69 次
- Moment Quantization for Video Temporal GroundingXiaolong Sun, Le Wang, Sanping Zhou, Liushuai Shi 等ICCV 2025 · 被引用 2 次
- Efficient Spatio-Temporal Video Grounding with Semantic-Guided Feature DecompositionWeikang Wang, Jing Liu, Yuting Su, Weizhi NieACM MM 2023 · 被引用 8 次
- CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal GroundingZhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao 等ACL 2023 · 被引用 17 次
