Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal Learning
Minghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng, Yang Liu
摘要
Temporal sentence grounding aims to detect the most salient moment corresponding to the natural language query from untrimmed videos. As labeling the temporal boundaries is labor-intensive and subjective, the weakly- supervised methods have recently received increasing attention. Most of the existing weakly-supervised methods gen-erate the proposals by sliding windows, which are content- independent and of low quality. Moreover, they train their model to distinguish positive visual-language pairs from negative ones randomly collected from other videos, ignoring the highly confusing video segments within the same video. In this paper, we propose Contrastive Proposal Learning(CPL) to overcome the above limitations. Specifi-cally, we use multiple learnable Gaussian functions to gen-erate both positive and negative proposals within the same video that can characterize the multiple events in a long video. Then, we propose a controllable easy to hard neg-ative proposal mining strategy to collect negative samples within the same video, which can ease the model opti-mization and enables CPL to distinguish highly confusing scenes. The experiments show that our method achieves state-of-the-art performance on Charades-STA and Activi-tyNet Captions datasets. The code and models are available at https://github.com/minghangz/cpl.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 被引用 44 次
- Phrase-Level Temporal Relationship Mining for Temporal Sentence LocalizationMinghang Zheng, Sizhe Li, Qingchao Chen, Yuxin Peng 等AAAI 2023 · 被引用 26 次
- Curriculum Multi-Negative Augmentation for Debiased Video GroundingXiaohan Lan, Yitian Yuan, Hong Chen, Xin Wang 等AAAI 2023 · 被引用 26 次
- Faster Video Moment Retrieval with Point-Level SupervisionXun Jiang, Zailei Zhou, Xing Xu, Yang Yang 等ACM MM 2023 · 被引用 24 次
- SCANet: Scene Complexity Aware Network for Weakly-Supervised Video Moment RetrievalSunjae Yoon, Gwanhyeong Koo, Dahyun Kim, Chang D. YooICCV 2023 · 被引用 24 次
它引用的顶会 Paper11
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang 等AAAI 2020 · 被引用 170 次
- Weakly Supervised Video Moment Localization with Contrastive Negative Sample MiningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yang LiuAAAI 2022 · 被引用 109 次
- Cross-Sentence Temporal and Semantic Relations in Video Activity LocalisationJiabo Huang, Yang Liu, Shaogang Gong, Hailin JinICCV 2021 · 被引用 77 次
- Visual Co-Occurrence Alignment Learning for Weakly-Supervised Video Moment RetrievalZheng Wang, Jingjing Chen, Yu-Gang JiangACM MM 2021 · 被引用 74 次
- Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed VideosJie Wu, Guanbin Li, Xiaoguang Han, Liang LinACM MM 2020 · 被引用 70 次
相关 Paper
- D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance AnnotationHanjun Li, Xiujun Shu, Sunan He, Ruizhi Qiao 等ICCV 2023 · 被引用 21 次
- Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence GroundingJiaming Chen, Weixin Luo, Wei Zhang, Lin MaAAAI 2022 · 被引用 33 次
- Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence LocalizationMinghang Zheng, Shaogang Gong, Hailin Jin, Yuxin Peng 等ACL 2023 · 被引用 18 次
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di 等AAAI 2022 · 被引用 52 次
- Weakly Supervised Temporal Sentence Grounding with Uncertainty-Guided Self-trainingYifei Huang, Lijin Yang, Yoichi SatoCVPR 2023
