Local-Global Video-Text Interactions for Temporal Grounding
Jonghwan Mun, Minsu Cho, Bohyung Han
Abstract
This paper addresses the problem of text-to-video temporal grounding, which aims to identify the time interval in a video semantically relevant to a text query. We tackle this problem using a novel regression-based model that learns to extract a collection of mid-level features for semantic phrases in a text query, which corresponds to important semantic entities described in the query (e.g., actors, objects, and actions), and reflect bi-modal interactions between the linguistic features of the query and the visual features of the video in multiple levels. The proposed method effectively predicts the target time interval by exploiting contextual information from local to global during bi-modal interactions. Through in-depth ablation studies, we find out that incorporating both local and global context in video and text interactions is crucial to the accurate grounding. Our experiment shows that the proposed method outperforms the state of the arts on Charades-STA and ActivityNet Captions datasets by large margins, 7.44% and 4.61% points at Recall@tIoU=0.5 metric, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32972d83-9406-4166-b69e-7774457b7c01Cited by top-tier papers126
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick et al.ICCV 2023 · 221 citations
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji et al.AAAI 2021 · 186 citations
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningSimon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxNeurIPS 2020 · 186 citations
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin et al.ICCV 2023 · 152 citations
Builds on3
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Dynamic Graph Attention for Referring Expression ComprehensionSibei Yang, Guanbin Li, Yizhou YuICCV 2019 · 251 citations
Related papers
- End-to-end Multi-modal Video Temporal GroundingYi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan YangNeurIPS 2021 · 68 citations
- Cascaded Prediction Network via Segment Tree for Temporal Video GroundingYang Zhao, Zhou Zhao, Zhu Zhang, Zhijie LinCVPR 2021
- Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou et al.CVPR 2021
- Phrase-Level Temporal Relationship Mining for Temporal Sentence LocalizationMinghang Zheng, Sizhe Li, Qingchao Chen, Yuxin Peng et al.AAAI 2023 · 26 citations
- Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding AbilityChenzhaoyu, Hongnan Lin, Yongwei Nie, Fei Ma et al.ICLR 2026 · 3 citations
