VideoTrack: Learning to Track Objects via Video Transformer
Fei Xie, Lei Chu, Jiahao Li, Yan Lu, Chao Ma
Abstract
Existing Siamese tracking methods, which are built on pair-wise matching between two single frames, heavily rely on additional sophisticated mechanism to exploit temporal information among successive video frames, hindering them from efficiency and industrial deployments. In this work, we resort to sequence-level target matching that can encode temporal contexts into the spatial features through a neat feedforward video model. Specifically, we adapt the standard video transformer architecture to visual tracking by enabling spatiotemporal feature learning directly from frame-level patch sequences. To better adapt to the tracking task, we carefully blend the spatiotemporal information in the video clips through sequential multi-branch triplet blocks, which formulates a video transformer backbone. Our experimental study compares different model variants, such as tokenization strategies, hierarchical structures, and video attention schemes. Then, we propose a disentangled dual-template mechanism that decouples static and dynamic appearance clues over time, and reduces temporal redundancy in video frames. Extensive experiments show that our method, named as VideoTrack, achieves state-ofthe-art results while running in real-time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2600e346-a067-4ac7-be2e-b1e905e8b7c8Cited by top-tier papers22
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.AAAI 2024 · 247 citations
- Explicit Visual Prompts for Visual Object TrackingLiangtao Shi, Bineng Zhong, Qihua Liang, Ning Li et al.AAAI 2024 · 117 citations
- Autoregressive Queries for Adaptive Tracking with Spatio-Temporal TransformersJinxia Xie, Bineng Zhong, Zhiyi Mo, Shengping Zhang et al.CVPR 2024 · 100 citations
- Exploring Enhanced Contextual Information for Video-Level Object TrackingBen Kang, Xin Chen, Simiao Lai, Yang Liu et al.AAAI 2025 · 48 citations
- Decoupled Spatio-Temporal Consistency Learning for Self-Supervised TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Ning Li et al.AAAI 2025 · 41 citations
Builds on33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
Related papers
- Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual TrackingNing Wang, Wengang Zhou, Jie Wang, Houqiang LiCVPR 2021
- SeqTrack: Sequence to Sequence Learning for Visual Object TrackingXin Chen, Houwen Peng, Dong Wang, Huchuan Lu et al.CVPR 2023
- An Efficient Token Compression Framework for Visual Object TrackingWeijing Wu, Qihua Liang, Bineng Zhong, Haiying Xia et al.CVPR 2026 · 1 citation
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 16 citations
- Fast Spatial Tracking with Visual Geometry TransformerChengjie Huang, GUILE WU, Dongfeng Bai, Bingbing LiuCVPR 2026
