Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence Grounding
Jiaming Chen, Weixin Luo, Wei Zhang, Lin Ma
Abstract
Weakly supervised temporal sentence grounding aims to temporally localize the target segment corresponding to a given natural language query, where it provides video-query pairs without temporal annotations during training. Most existing methods use the fused visual-linguistic feature to reconstruct the query, where the least reconstruction error determines the target segment. This work introduces a novel approach that explores the inter-contrast between videos in a composed video by selecting components from two different videos and fusing them into a single video. Such a straightforward yet effective composition strategy provides the temporal annotations at multiple composed positions, resulting in numerous videos with temporal ground-truths for training the temporal sentence grounding task. A transformer framework is introduced with multi-tasks training to learn a compact but efficient visual-linguistic space. The experimental results on the public Charades-STA and ActivityNet-Caption dataset demonstrate the effectiveness of the proposed method, where our approach achieves comparable performance over the state-of-the-art weakly-supervised baselines. The code is available at https://github.com/PPjmchen/ Composition WSTG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b458baa1-577b-401c-9d30-497271925a65Cited by top-tier papers9
- Hypotheses Tree Building for One-Shot Temporal Sentence LocalizationDaizong Liu, Xiang Fang, Pan Zhou, Xing Di et al.AAAI 2023 · 29 citations
- Faster Video Moment Retrieval with Point-Level SupervisionXun Jiang, Zailei Zhou, Xing Xu, Yang Yang et al.ACM MM 2023 · 24 citations
- Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video GroundingSunoh Kim, Jungchan Cho, Joonsang Yu, Youngjoon Yoo et al.AAAI 2024 · 19 citations
- Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using LanguageXiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu et al.ACM MM 2024 · 8 citations
- Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingChaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang HuCVPR 2024 · 5 citations
Builds on6
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang et al.AAAI 2020 · 170 citations
- Adaptive Reconstruction Network for Weakly Supervised Referring Expression GroundingXuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha et al.ICCV 2019 · 93 citations
- Regularized Two-Branch Proposal Networks for Weakly-Supervised Moment Retrieval in VideosZhu Zhang, Zhijie Lin, Zhou Zhao, Jieming Zhu et al.ACM MM 2020 · 86 citations
- Detector-Free Weakly Supervised Grounding by SeparationAssaf Arbelle, Sivan Doveh, Amit Alfassy, Joseph Shtok et al.ICCV 2021 · 31 citations
Related papers
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng et al.CVPR 2022 · 108 citations
- DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-to-Fine Contrastive RankingLijin Yang, Quan Kong, Hsuan-Kung Yang, Wadim Kehl et al.CVPR 2023
- Weakly Supervised Temporal Sentence Grounding with Uncertainty-Guided Self-trainingYifei Huang, Lijin Yang, Yoichi SatoCVPR 2023
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di et al.AAAI 2022 · 52 citations
