Semi-supervised Video Paragraph Grounding with Contrastive Encoder
Xun Jiang, Xing Xu, Jingran Zhang, Fumin Shen, Zuo Cao, Heng Tao Shen
Abstract
Video events grounding aims at retrieving the most relevant moments from an untrimmed video in terms of a given natural language query. Most previous works focus on Video Sentence Grounding (VSG), which localizes the moment with a sentence query. Recently, researchers extended this task to Video Paragraph Grounding (VPG) by retrieving multiple events with a paragraph. However, we find the existing VPG methods may not perform well on context modeling and highly rely on video-paragraph annotations. To tackle this problem, we propose a novel VPG method termed Semi-supervised Video-Paragraph TRansformer (SVPTR), which can more effectively exploit contextual information in paragraphs and significantly reduce the dependency on annotated data. Our SVPTR method consists of two key components: (1) a base model VPTR that learns the videoparagraph alignment with contrastive encoders and tackles the lack of sentence-level contextual interactions and (2) a semi-supervised learning framework with multimodal feature perturbations that reduces the requirements of annotated training data. We evaluate our model on three widelyused video grounding datasets, i.e., ActivityNet-Caption, Charades-CD-OOD, and TACoS. The experimental results show that our SVPTR method establishes the new state-ofthe-art performance on all datasets. Even under the conditions of fewer annotations, it can also achieve competitive results compared with recent VPG methods. * Corresponding author. (b) Video Paragraph Grounding Two young girls are standing in the kitchen preparing to cook. They then open a box of brownies…get an egg out of the fridge. After, the two continue to stir the contents … placing them on a pan. Once the cookies are … watch the cookies bake. When they are done, they … begin eating. 15.67s 28.73s 73.78s 105.12s 130.59s 0s Multi-Multi Localization Sentence: The man with red shorts serves the ball. (a) Video Sentence Grounding 12.91s 13.63s
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c01f0877-c974-439c-8895-3a5b770d9bf1Cited by top-tier papers14
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen et al.AAAI 2024 · 59 citations
- DePT: Decoupled Prompt TuningJi Zhang, Shihan Wu, Lianli Gao, Heng Tao Shen et al.CVPR 2024 · 36 citations
- Faster Video Moment Retrieval with Point-Level SupervisionXun Jiang, Zailei Zhou, Xing Xu, Yang Yang et al.ACM MM 2023 · 24 citations
- DQS3D: Densely-matched Quantization-aware Semi-supervised 3D DetectionHuan-ang Gao, Beiwen Tian, Pengfei Li, Hao Zhao et al.ICCV 2023 · 21 citations
- Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingChaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang HuCVPR 2024 · 5 citations
Builds on17
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 206 citations
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji et al.AAAI 2021 · 186 citations
- Proposal-Free Video Grounding with Contextual Pyramid NetworkKun Li, Dan Guo, Meng WangAAAI 2021 · 138 citations
- Tree-Structured Policy Based Progressive Reinforcement Learning for Temporally Language Grounding in VideoJie Wu, Guanbin Li, Si Liu, Liang LinAAAI 2020 · 117 citations
Related papers
- Hierarchical Semantic Correspondence Networks for Video Paragraph GroundingChaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng et al.CVPR 2023
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi et al.ACM MM 2024 · 2 citations
- Dense Events Grounding in VideoPeijun Bao, Qian Zheng, Yadong MuAAAI 2021 · 37 citations
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 75 citations
- On Pursuit of Designing Multi-modal Transformer for Video GroundingMeng Cao, Long Chen, Mike Zheng Shou, Can Zhang et al.EMNLP 2021 · 63 citations
