Correlation-Aware Deep Tracking
Fei Xie, Chunyu Wang, Guangting Wang, Yue Cao, Wankou Yang, Wenjun Zeng
Abstract
Robustness and discrimination power are two fundamental requirements in visual object tracking. In most tracking paradigms, we find that the features extracted by the popular Siamese-like networks cannot fully discriminatively model the tracked targets and distractor objects, hindering them from simultaneously meeting these two requirements. While most methods focus on designing robust correlation operations, we propose a novel target-dependent feature network inspired by the self-/cross-attention scheme. In contrast to the Siamese-like feature extraction, our network deeply embeds cross-image feature correlation in multiple layers of the feature network. By extensively matching the features of the two images through multiple layers, it is able to suppress non-target features, resulting in instancevarying feature extraction. The output features of the search image can be directly used for predicting target locations without extra correlation step. Moreover, our model can be flexibly pre-trained on abundant unpaired images, leading to notably faster convergence than the existing methods. Extensive experiments show our method achieves the stateof-the-art results while running at real-time. Our feature networks also can be applied to existing tracking pipelines seamlessly to raise the tracking performance. 𝑓 𝑧 z (a1) Siamese-like feature network x Correlation Operation (b1) Feature correlation (a2) Target-dependent feature network 𝑓 𝑧 𝑓 𝑥 (b2) Ours z x 𝑓 cor 𝑓 𝑥 𝑓 cor (c) Prediction Localization Size estimation Prediction head 𝑓 cor
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers42
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu et al.NeurIPS 2022 · 556 citations
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.AAAI 2024 · 247 citations
- MixFormerV2: Efficient Fully Transformer TrackingYutao Cui, Tianhui Song, Gangshan Wu, Limin WangNeurIPS 2023 · 193 citations
- Robust Object Modeling for Visual TrackingYidong Cai, Jie Liu, Jie Tang, Gangshan WuICCV 2023 · 165 citations
- Exploring Lightweight Hierarchical Vision Transformers for Efficient Visual TrackingBen Kang, Xin Chen, Dong Wang, Houwen Peng et al.ICCV 2023 · 119 citations
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang et al.NeurIPS 2021 · 1,388 citations
Related papers
- Reinforced Similarity Learning: Siamese Relation Networks for Robust Object TrackingDawei Zhang, Zhonglong Zheng, Minglu Li, Xiaowei He et al.ACM MM 2020 · 15 citations
- Deformable Siamese Attention Networks for Visual Object TrackingYuechen Yu, Yilei Xiong, Weilin Huang, Matthew R. ScottCVPR 2020
- Discriminative and Robust Online Learning for Siamese Visual TrackingJinghao Zhou, Peng Wang, Haoyang SunAAAI 2020 · 66 citations
- Graph Attention TrackingDongyan Guo, Yanyan Shao, Ying Cui, Zhenhua Wang et al.CVPR 2021
- Deep Meta Learning for Real-Time Target-Aware Visual TrackingJanghoon Choi, Junseok Kwon, Kyoung Mu LeeICCV 2019 · 93 citations
