Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
Xin Gu, Yaojie Shen, Chenxi Luo, Tiejian Luo, Yan Huang, Yuewei Lin, Heng Fan, Libo Zhang
摘要
Transformer has attracted increasing interest in spatio-temporal video grounding, or STVG, owing to its end-to-end pipeline and promising result. Existing Transformerbased STVG approaches often leverage a set of object queries, which are initialized simply using zeros and then gradually learn target position information via iterative interactions with multimodal features, for spatial and temporal localization. Despite simplicity, these zero object queries, due to lacking target-specific cues, are hard to learn discriminative target information from interactions with multimodal features in complicated scenarios (e.g., with distractors or occlusion), resulting in degradation. Addressing this, we introduce a novel Target-Aware Transformer for STVG (TA-STVG), which seeks to adaptively generate object queries via exploring targetspecific cues from the given video-text pair, for improving STVG. The key lies in two simple yet effective modules, comprising text-guided temporal sampling (TTS) and attribute-aware spatial activation (ASA), working in a cascade. The former focuses on selecting target-relevant temporal cues from a video utilizing holistic text information, while the latter aims at further exploiting the fine-grained visual attribute information of the object from previous target-aware temporal cues, which is applied for object query initialization. Compared to existing methods leveraging zero-initialized queries, object queries in our TA-STVG, directly generated from a given video-text pair, naturally carry target-specific cues, making them adaptive and better interact with multimodal features for learning more discriminative information to improve STVG. In our experiments on three benchmarks, including HCSTVG-v1/-v2 and VidSTG, TA-STVG achieves state-of-the-art performance and significantly outperforms the baseline, validating its efficacy. Moreover, TTS and ASA are designed for general purpose. When applied to existing methods such as TubeDETR and STCAT, we show substantial performance gains, verifying its generality. Code is released at https://github.com/HengLan/TA-STVG.
Recently, inspired by the compact end-to-end pipeline of Transformer-based detection (Carion et al., 2020), researchers have introduced the Transformer (Vaswani et al., 2017) for STVG given that they both are localization task, and achieved previously unattainable result
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- VGR: Visual Grounded ReasoningJiacong Wang, Zijian Kang, Haochen Wang, Xiao Liang 等ICLR 2026 · 被引用 64 次
- OmniSTVG: Toward Spatio-Temporal Omni-Object Video GroundingJiali Yao, Xin Gu, Xinran Deng, Mengrui Dai 等ICLR 2026 · 被引用 8 次
- OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex ScenariosHong Gao, Jingyu Wu, Xiangkai Xu, Kangni Xie 等CVPR 2026 · 被引用 4 次
- STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement LearningXiaowen Zhang, Zhi Gao, Licheng Jiao, Lingling Li 等ICLR 2026 · 被引用 3 次
- SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video GroundingHong Gao, Xiangkai Xu, Bin Zhong, Junjie Yin 等CVPR 2026
它引用的顶会 Paper39
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
相关 Paper
- Context-Guided Spatio-Temporal Video GroundingXin Gu, Heng Fan, Yan Huang, Tiejian Luo 等CVPR 2024
- TubeDETR: Spatio-Temporal Video Grounding with TransformersAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等CVPR 2022 · 被引用 87 次
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 被引用 75 次
- Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video GroundingYang Jin, Yongzhi Li, Zehuan Yuan, Yadong MuNeurIPS 2022 · 被引用 69 次
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video GroundingZaiquan Yang, Yuhao Liu, Gerhard P. Hancke, Rynson W. H. LauNeurIPS 2025 · 被引用 10 次
