Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos
Yulin Pan, Xiangteng He, Biao Gong, Yiliang Lv, Yujun Shen, Yuxin Peng, Deli Zhao
Abstract
Video temporal grounding aims to pinpoint a video segment that matches the query description. Despite the recent advance in short-form videos (e.g., in minutes), temporal grounding in long videos (e.g., in hours) is still at its early stage. To address this challenge, a common practice is to employ a sliding window, yet can be inefficient and inflexible due to the limited number of frames within the window. In this work, we propose an end-to-end framework for fast temporal grounding, which is able to model an hours-long video with one-time network execution. Our pipeline is formulated in a coarse-to-fine manner, where we first extract context knowledge from non-overlapped video clips (i.e., anchors), and then supplement the anchors that highly response to the query with detailed content knowledge. Besides the remarkably high pipeline efficiency, another advantage of our approach is the capability of capturing long-range temporal correlation, thanks to modeling the entire video as a whole, and hence facilitates more accurate grounding. Experimental results suggest that, on the long-form video datasets MAD and Ego4d, our method significantly outperforms state-ofthe-arts, and achieves 14.6× / 102.8× higher efficiency respectively. Project can be found at https://github. com/afcedf/SOONet.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bc3c240-a911-4b52-8689-9e44ca5f2475Cited by top-tier papers19
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang et al.NeurIPS 2025 · 30 citations
- SnAG: Scalable and Accurate Video GroundingFangzhou Mu, Sicheng Mo, Yin LiCVPR 2024 · 13 citations
- Moment Detection in Long Tutorial VideosIoana Croitoru, Simion-Vlad Bogolin, Samuel Albanie, Yang Liu et al.ICCV 2023 · 7 citations
- Temporal Sentence Grounding in Streaming VideosTian Gan, Xiao Wang, Yan Sun, Jianlong Wu et al.ACM MM 2023 · 5 citations
- OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal GroundingMinghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng et al.CVPR 2026 · 4 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- Localizing Moments in Long Video Via Multimodal GuidanceWayner Barrios, Mattia Soldan, Alberto Mario Ceballos-Arroyo, Fabian Caba Heilbron et al.ICCV 2023 · 32 citations
- CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal GroundingZhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao et al.ACL 2023 · 17 citations
- Hierarchical Event Memory for Accurate and Low-Latency Online Video Temporal GroundingMinghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang et al.ICCV 2025 · 3 citations
- On Pursuit of Designing Multi-modal Transformer for Video GroundingMeng Cao, Long Chen, Mike Zheng Shou, Can Zhang et al.EMNLP 2021 · 63 citations
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 206 citations
