Efficient Spatio-Temporal Video Grounding with Semantic-Guided Feature Decomposition
Weikang Wang, Jing Liu, Yuting Su, Weizhi Nie
Abstract
Spatio-temporal video grounding (STVG) aims to localize the spatio-temporal object tube in a video according to a given text query. Current approaches address the STVG task with end-to-end frameworks while suffering from heavy computational complexity and insufficient spatio-temporal interactions. To overcome these limitations, we propose a novel Semantic-Guided Feature Decomposition based Network (SGFDN). A semantic-guided mapping operation is proposed to decompose the 3D spatio-temporal feature into 2D motions and 1D object embedding without losing much object-related semantic information. Thus, the computational complexity in computationally expensive operations such as attention mechanisms can be effectively reduced by replacing the input spatio-temporal feature with the decomposed features. Furthermore, based on this decomposition strategy, a pyramid relevance filtering based attention is proposed to capture the cross-modal interactions at multiple spatio-temporal scales. In addition, a decomposition-based grounding head is proposed to locate the queried objects with less computational complexity. Extensive experiments on two widely-used STVG datasets (VidSTG and HC-STVG) demonstrate that our method enjoys state-of-the-art performance as well as less computational complexity. The code has been available at https://github.com/TJUMMG/SGFDN.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Where Does It Exist from the Low-Altitude: Spatial Aerial Video GroundingYang Zhan, Yuan YuanNeurIPS 2025 · 8 citations
- OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex ScenariosHong Gao, Jingyu Wu, Xiangkai Xu, Kangni Xie et al.CVPR 2026 · 4 citations
- TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video GroundingJinxuan Li, Yi Zhang, Jian-Fang Hu, Chaolei Tan et al.AAAI 2026 · 1 citation
- Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video GroundingXin Gu, Yaojie Shen, Chenxi Luo, Tiejian Luo et al.ICLR 2025
Related papers
- Context-Guided Spatio-Temporal Video GroundingXin Gu, Heng Fan, Yan Huang, Tiejian Luo et al.CVPR 2024
- Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form SentencesZhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang et al.CVPR 2020
- Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video GroundingYang Jin, Yongzhi Li, Zehuan Yuan, Yadong MuNeurIPS 2022 · 69 citations
- TubeDETR: Spatio-Temporal Video Grounding with TransformersAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.CVPR 2022 · 87 citations
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 75 citations
