Exploring Reliable Spatiotemporal Dependencies for Efficient Visual Tracking
Junze Shi, Yang Yu, Jian Shi, Haibo Luo
Abstract
Recent advances in transformer-based lightweight object tracking have established new standards across benchmarks, leveraging the global receptive field and powerful feature extraction capabilities of attention mechanisms. Despite these achievements, existing methods universally employ sparse sampling during training—utilizing only one template and one search image per sequence—which fails to comprehensively explore spatiotemporal information in videos. This limitation constrains performance and causes the gap between lightweight and high-performance trackers. To bridge this divide while maintaining real-time efficiency, we propose STDTrack, a framework that pioneers the integration of reliable spatiotemporal dependencies into lightweight trackers. Our approach implements dense video sampling to maximize spatiotemporal information utilization. We introduce a temporally propagating spatiotemporal token to guide per-frame feature extraction. To ensure comprehensive target state representation, we design the Multi-frame Information Fusion Module (MFIFM), which augments current dependencies using historical context. The MFIFM operates on features stored in our constructed Spatiotemporal Token Maintainer (STM), where a quality-based update mechanism ensures information reliability. Considering the scale variation among tracking targets, we develop a multi-scale prediction head to dynamically adapt to objects of different sizes. Extensive experiments demonstrate state-of-the-art results across six benchmarks. Notably, on GOT-10k, STDTrack rivals certain high-performance non-real-time trackers (e.g., MixFormer) while operating at 192 FPS (GPU) and 41 FPS (CPU).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu et al.NeurIPS 2022 · 556 citations
- Transforming Model Prediction for TrackingChristoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul et al.CVPR 2022 · 399 citations
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.AAAI 2024 · 247 citations
Related papers
- Explicit Visual Prompts for Visual Object TrackingLiangtao Shi, Bineng Zhong, Qihua Liang, Ning Li et al.AAAI 2024 · 117 citations
- Less Is More: Token Context-Aware Learning for Object TrackingChenlong Xu, Bineng Zhong, Qihua Liang, Yaozong Zheng et al.AAAI 2025
- Drift-Resilient Temporal Priors for Visual TrackingYuqing Huang, Liting Lin, Weijun Zhuang, Zhenyu He et al.CVPR 2026 · 1 citation
- Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual TrackingNing Wang, Wengang Zhou, Jie Wang, Houqiang LiCVPR 2021
- Robust Object Modeling for Visual TrackingYidong Cai, Jie Liu, Jie Tang, Gangshan WuICCV 2023 · 165 citations
