Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding
Jongbhin Woo, Hyeonggon Ryu, Youngjoon Jang, Jae-Won Cho, Joon Son Chung
Abstract
Video Temporal Grounding (VTG) aims to identify visual frames in a video clip that match text queries. Recent studies in VTG employ cross-attention to correlate visual frames and text queries as individual token sequences. However, these approaches overlook a crucial aspect of the problem: a holistic understanding of the query sentence. A model may capture correlations between individual word tokens and arbitrary visual frames while possibly missing out on the global meaning. To address this, we introduce two primary contributions: (1) a visual frame-level gate mechanism that incorporates holistic textual information, (2) cross-modal alignment loss to learn the fine-grained correlation between query and relevant frames. As a result, we regularize the effect of individual word tokens and suppress irrelevant visual frames. We demonstrate that our method outperforms state-of-the-art approaches in VTG benchmarks, indicating that holistic text understanding guides the model to focus on the semantically important parts within the video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3aa4fb27-060f-48e7-b856-54962f1b1f67Cited by top-tier papers2
- Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal GroundingMinseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim et al.NeurIPS 2025 · 4 citations
- Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video CaptioningSeungHyup Baek, Jimin Lee, Hyeongkeun Lee, Jae Won ChoCVPR 2026 · 1 citation
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
Related papers
- T2SGrid: Temporal-to-Spatial Gridification for Video Temporal GroundingChaohong Guo, Yihan He, Yongwei Nie, Fei Ma et al.CVPR 2026 · 2 citations
- Look Around Before Locating: Considering Content and Structure Information for Visual GroundingShiyi Zheng, Peizhi Zhao, Zhilong Zheng, Peihang He et al.AAAI 2025 · 3 citations
- Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Pan ZhouEMNLP 2021 · 34 citations
- Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer NetworkXiang Fang, Wanlong Fang, Changshuo Wang, Daizong Liu et al.AAAI 2025 · 10 citations
- Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video GroundingYang Jin, Yongzhi Li, Zehuan Yuan, Yadong MuNeurIPS 2022 · 69 citations
