Siam R-CNN: Visual Tracking by Re-Detection
Paul Voigtlaender, Jonathon Luiten, Philip H. S. Torr, Bastian Leibe
Abstract
We present Siam R-CNN, a Siamese re-detection architecture which unleashes the full power of two-stage object detection approaches for visual object tracking. We combine this with a novel tracklet-based dynamic programming algorithm, which takes advantage of re-detections of both the first-frame template and previous-frame predictions, to model the full history of both the object to be tracked and potential distractor objects. This enables our approach to make better tracking decisions, as well as to re-detect tracked objects after long occlusion. Finally, we propose a novel hard example mining strategy to improve Siam R-CNN's robustness to similar looking objects. Siam R-CNN achieves the current best performance on ten tracking benchmarks, with especially strong results for long-term tracking. We make our code and models available at www. vision.rwth-aachen.de/page/siamrcnn.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers63
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu et al.NeurIPS 2022 · 556 citations
- Transformer Tracking with Cyclic Shifting Window AttentionZikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei YangCVPR 2022 · 220 citations
- Compact Transformer Tracker with Correlative Masked ModelingZikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen et al.AAAI 2023 · 136 citations
- Divert More Attention to Vision-Language TrackingMingzhe Guo, Zhipeng Zhang, Heng Fan, Liping JingNeurIPS 2022 · 122 citations
Builds on5
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- 'Skimming-Perusal' Tracking: A Framework for Real-Time and Robust Long-Term TrackingBin Yan, Haojie Zhao, Dong Wang, Huchuan Lu et al.ICCV 2019 · 177 citations
- Bridging the Gap Between Detection and Tracking: A Unified ApproachLianghua Huang, Xin Zhao, Kaiqi HuangICCV 2019 · 40 citations
Related papers
- Discriminative and Robust Online Learning for Siamese Visual TrackingJinghao Zhou, Peng Wang, Haoyang SunAAAI 2020 · 66 citations
- SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual TrackingDongyan Guo, Jun Wang, Ying Cui, Zhenhua Wang et al.CVPR 2020
- Reinforced Similarity Learning: Siamese Relation Networks for Robust Object TrackingDawei Zhang, Zhonglong Zheng, Minglu Li, Xiaowei He et al.ACM MM 2020 · 15 citations
- Learning To Filter: Siamese Relation Network for Robust TrackingSiyuan Cheng, Bineng Zhong, Guorong Li, Xin Liu et al.CVPR 2021
- Learning the Model Update for Siamese TrackersLichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer, Martin Danelljan et al.ICCV 2019 · 371 citations
