Learn to Match: Automatic Matching Network Design for Visual Tracking
Zhipeng Zhang, Yihao Liu, Xiao Wang, Bing Li, Weiming Hu
Abstract
Siamese tracking has achieved groundbreaking performance in recent years, where the essence is the efficient matching operator cross-correlation and its variants. Besides the remarkable success, it is important to note that the heuristic matching network design relies heavily on expert experience. Moreover, we experimentally find that one sole matching operator is difficult to guarantee stable tracking in all challenging environments. Thus, in this work, we introduce six novel matching operators from the perspective of feature fusion instead of explicit similarity learning , namely Concatenation, Pointwise-Addition, Pairwise-Relation, FiLM, Simple-Transformer and Transductive-Guidance, to explore more feasibility on matching operator selection. The analyses reveal these operators' selective adaptability on different environment degradation types, which inspires us to combine them to explore complementary features. To this end, we propose binary channel manipulation (BCM) to search for the optimal combination of these operators. BCM determines to retrain or discard one operator by learning its contribution to other tracking steps. By inserting the learned matching networks to a strong baseline tracker Ocean [47], our model achieves favorable gains by 67 . 2 → 71 . 4 , 52 . 6 → 58 . 3 , 70 . 3 → 76 . 0 success on OTB100, LaSOT, and TrackingNet, respectively. Notably, Our tracker, dubbed AutoMatch , uses less than half of training data/time than the baseline tracker, and runs at 50 FPS using PyTorch. Code and model are released at https://github.com/JudasDie/SOTS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2891a05b-d89e-4e66-a328-5871a9400b0bCited by top-tier papers31
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu et al.NeurIPS 2022 · 556 citations
- Transformer Tracking with Cyclic Shifting Window AttentionZikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei YangCVPR 2022 · 220 citations
- Correlation-Aware Deep TrackingFei Xie, Chunyu Wang, Guangting Wang, Yue Cao et al.CVPR 2022 · 189 citations
- Compact Transformer Tracker with Correlative Masked ModelingZikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen et al.AAAI 2023 · 136 citations
Builds on12
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan et al.AAAI 2020 · 944 citations
- GlobalTrack: A Simple and Strong Baseline for Long-Term TrackingLianghua Huang, Xin Zhao, Kaiqi HuangAAAI 2020 · 278 citations
- GradNet: Gradient-Guided Network for Visual Object TrackingPeixia Li, Boyu Chen, Wanli Ouyang, Dong Wang et al.ICCV 2019 · 255 citations
- Batch-shaping for learning conditional channel gated networksBabak Ehteshami Bejnordi, Tijmen Blankevoort, Max WellingICLR 2020 · 82 citations
Related papers
- Deformable Siamese Attention Networks for Visual Object TrackingYuechen Yu, Yilei Xiong, Weilin Huang, Matthew R. ScottCVPR 2020
- Reinforced Similarity Learning: Siamese Relation Networks for Robust Object TrackingDawei Zhang, Zhonglong Zheng, Minglu Li, Xiaowei He et al.ACM MM 2020 · 15 citations
- Transformer TrackingXin Chen, Bin Yan, Jiawen Zhu, Dong Wang et al.CVPR 2021
- HIPTrack: Visual Tracking with Historical PromptsWenrui Cai, Qingjie Liu, Yunhong WangCVPR 2024
- Learning To Fuse Asymmetric Feature Maps in Siamese TrackersWencheng Han, Xingping Dong, Fahad Shahbaz Khan, Ling Shao et al.CVPR 2021
