Dense Events Grounding in Video
Peijun Bao, Qian Zheng, Yadong Mu
Abstract
This paper explores a novel setting of temporal sentence grounding for the first time, dubbed as dense events grounding. Given an untrimmed video and a paragraph description, dense events grounding aims to jointly localize temporal moments of multiple events described in the paragraph. Our main motivating fact is that multiple events to be grounded in a video are often semantically related and temporally coordinated according to their order appearing in the paragraph. This fact sheds light on devising more accurate visual grounding model. In this work, we propose Dense Events Propagation Network (DepNet) for this novel task. DepNet first adaptively aggregates temporal and semantic information of dense events into a compact set through a second-order attention pooling, then selectively propagates the aggregated information to each single event with soft attention. Based on such aggregation-and-propagation mechanism, DepNet can effectively exploit both the temporal order and semantic relations of dense events. We conduct comprehensive experiments on large-scale datasets ActivityNet Captions and TACoS. For fair comparisons, our evaluations include both state-of-art single-event grounding methods and their natural extensions to the dense-events grounding setting implemented by us. All experiments clearly shows the performance superiority of the proposed DepNet by significant margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1699ecc-2635-4a0b-8b16-eaf856fdce60Cited by top-tier papers15
- Semi-supervised Video Paragraph Grounding with Contrastive EncoderXun Jiang, Xing Xu, Jingran Zhang, Fumin Shen et al.CVPR 2022 · 42 citations
- Learning to Ground Instructional Articles in Videos through NarrationsEffrosyni Mavroudi, Triantafyllos Afouras, Lorenzo TorresaniICCV 2023 · 28 citations
- Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video GroundingPeijun Bao, Yong Xia, Wenhan Yang, Boon Poh Ng et al.AAAI 2024 · 20 citations
- Dense Object Grounding in 3D ScenesWencan Huang, Daizong Liu, Wei HuACM MM 2023 · 16 citations
- Weakly-Supervised Temporal Article GroundingLong Chen, Yulei Niu, Brian Chen, Xudong Lin et al.EMNLP 2022 · 11 citations
Builds on4
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 206 citations
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao et al.AAAI 2020 · 182 citations
Related papers
- Hierarchical Semantic Correspondence Networks for Video Paragraph GroundingChaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng et al.CVPR 2023
- Sketch, Ground, and Refine: Top-Down Dense Video CaptioningChaorui Deng, Shizhe Chen, Da Chen, Yuan He et al.CVPR 2021
- Cascaded Prediction Network via Segment Tree for Temporal Video GroundingYang Zhao, Zhou Zhao, Zhu Zhang, Zhijie LinCVPR 2021
- Exploiting Auxiliary Caption for Video GroundingHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li et al.AAAI 2024 · 16 citations
- Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou et al.CVPR 2021
