Learning Spatio-Temporal Transformer for Visual Tracking
Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, Huchuan Lu
Abstract
In this paper, we present a new tracking architecture with an encoder-decoder transformer as the key component. The encoder models the global spatio-temporal feature dependencies between target objects and search regions, while the decoder learns a query embedding to predict the spatial positions of the target objects. Our method casts object tracking as a direct bounding box prediction problem, without using any proposals or predefined anchors. With the encoder-decoder transformer, the prediction of objects just uses a simple fully-convolutional network, which estimates the corners of objects directly. The whole method is end-to-end, does not need any postprocessing steps such as cosine window and bounding box smoothing, thus largely simplifying existing tracking pipelines. The proposed tracker achieves state-of-the-art performance on multiple challenging short-term and long-term benchmarks, while running at real-time speed, being 6× faster than Siam R-CNN [54]. Code and models are open-sourced at https://github.com/researchmm/Stark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61739e41-035b-47e1-b8c1-cd53bb6b262dCited by top-tier papers139
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu et al.NeurIPS 2022 · 556 citations
- Transforming Model Prediction for TrackingChristoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul et al.CVPR 2022 · 399 citations
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein et al.ICCV 2023 · 255 citations
- TCTrack: Temporal Contexts for Aerial TrackingZiang Cao, Ziyuan Huang, Liang Pan, Shiwei Zhang et al.CVPR 2022 · 233 citations
Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen et al.NeurIPS 2020 · 2,118 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan et al.AAAI 2020 · 944 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
Related papers
- High-Performance Discriminative Tracking with TransformersBin Yu, Ming Tang, Linyu Zheng, Guibo Zhu et al.ICCV 2021 · 113 citations
- SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual TrackingDongyan Guo, Jun Wang, Ying Cui, Zhenhua Wang et al.CVPR 2020
- SeqTrack: Sequence to Sequence Learning for Visual Object TrackingXin Chen, Houwen Peng, Dong Wang, Huchuan Lu et al.CVPR 2023
- Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual TrackingNing Wang, Wengang Zhou, Jie Wang, Houqiang LiCVPR 2021
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 180 citations
