Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding
Zihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin, Tiancai Ye, Wei-Shi Zheng
Abstract
Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the language description, which requires effective joint modeling of spatio-temporal visual-linguistic dependencies. In this work, we propose a novel framework in which a static vision-language stream and a dynamic vision-language stream are developed to collaboratively reason the target tube. The static stream performs cross-modal understanding in a single frame and learns to attend to the target object spatially according to intraframe visual cues like object appearances. The dynamic stream models visual-linguistic dependencies across multiple consecutive frames to capture dynamic cues like motions. We further design a novel cross-stream collaborative block between the two streams, which enables the static and dynamic streams to transfer useful and complementary information from each other to achieve collaborative reasoning. Experimental results show the effectiveness of the collaboration of the two streams and our overall framework achieves new state-of-the-art performance on both HCSTVG and VidSTG datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilityJiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li et al.AAAI 2026 · 20 citations
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren et al.NeurIPS 2024 · 13 citations
- Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLMJunxiao Ma, Jingjing Wang, Jiamin Luo, Peiying Yu et al.WWW 2025 · 11 citations
- VideoGrounding-DINO: Towards Open-Vocabulary Spatio- Temporal Video GroundingSyed Talal Wasim, Muzammal Naseer, Salman H. Khan, Ming-Hsuan Yang et al.CVPR 2024 · 10 citations
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video GroundingZaiquan Yang, Yuhao Liu, Gerhard P. Hancke, Rynson W. H. LauNeurIPS 2025 · 10 citations
Builds on15
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji et al.AAAI 2021 · 186 citations
Related papers
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 75 citations
- Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form SentencesZhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang et al.CVPR 2020
- Agentic Spatio-Temporal Grounding via Collaborative ReasoningHeng Zhao, Yew-Soon Ong, Joey Tianyi ZhouSIGIR 2026 · 1 citation
- Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video GroundingYang Jin, Yongzhi Li, Zehuan Yuan, Yadong MuNeurIPS 2022 · 69 citations
- Context-Guided Spatio-Temporal Video GroundingXin Gu, Heng Fan, Yan Huang, Tiejian Luo et al.CVPR 2024
