Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual Tracking
Ning Wang, Wengang Zhou, Jie Wang, Houqiang Li
Abstract
In video object tracking, there exist rich temporal contexts among successive frames, which have been largely overlooked in existing trackers. In this work, we bridge the individual video frames and explore the temporal contexts across them via a transformer architecture for robust object tracking. Different from classic usage of the transformer in natural language processing tasks, we separate its encoder and decoder into two parallel branches and carefully design them within the Siamese-like tracking pipelines. The transformer encoder promotes the target templates via attentionbased feature reinforcement, which benefits the high-quality tracking model generation. The transformer decoder propagates the tracking cues from previous templates to the current frame, which facilitates the object searching process. Our transformer-assisted tracking framework is neat and trained in an end-to-end manner. With the proposed transformer, a simple Siamese matching approach is able to outperform the current top-performing trackers. By combining our transformer with the recent discriminative tracking pipeline, our method sets several new state-of-the-art records on prevalent tracking benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers88
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu et al.NeurIPS 2022 · 556 citations
- Focal Modulation NetworksJianwei Yang, Chunyuan Li, Xiyang Dai, Jianfeng GaoNeurIPS 2022 · 494 citations
- Transforming Model Prediction for TrackingChristoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul et al.CVPR 2022 · 399 citations
- Learning Target Candidate Association to Keep Track of What Not to TrackChristoph Mayer, Martin Danelljan, Danda Pani Paudel, Luc Van GoolICCV 2021 · 356 citations
Builds on10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan et al.AAAI 2020 · 944 citations
- Learning the Model Update for Siamese TrackersLichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer, Martin Danelljan et al.ICCV 2019 · 371 citations
- POST: POlicy-Based Switch TrackingNing Wang, Wengang Zhou, Guojun Qi, Houqiang LiAAAI 2020 · 9 citations
Related papers
- VideoTrack: Learning to Track Objects via Video TransformerFei Xie, Lei Chu, Jiahao Li, Yan Lu et al.CVPR 2023
- Target-Aware Tracking with Long-Term Context AttentionKaijie He, Canlong Zhang, Sheng Xie, Zhixin Li et al.AAAI 2023 · 102 citations
- High-Performance Discriminative Tracking with TransformersBin Yu, Ming Tang, Linyu Zheng, Guibo Zhu et al.ICCV 2021 · 113 citations
- SeqTrack: Sequence to Sequence Learning for Visual Object TrackingXin Chen, Houwen Peng, Dong Wang, Huchuan Lu et al.CVPR 2023
- TGTrack: Temporal Generative Learning for Unified Single Object TrackingWanting Geng, Xin Chen, Chuanyu Sun, Jie Zhao et al.CVPR 2026
