Correlation-Aware Deep Tracking
Fei Xie, Chunyu Wang, Guangting Wang, Yue Cao, Wankou Yang, Wenjun Zeng
摘要
Robustness and discrimination power are two fundamental requirements in visual object tracking. In most tracking paradigms, we find that the features extracted by the popular Siamese-like networks cannot fully discriminatively model the tracked targets and distractor objects, hindering them from simultaneously meeting these two requirements. While most methods focus on designing robust correlation operations, we propose a novel target-dependent feature network inspired by the self-/cross-attention scheme. In contrast to the Siamese-like feature extraction, our network deeply embeds cross-image feature correlation in multiple layers of the feature network. By extensively matching the features of the two images through multiple layers, it is able to suppress non-target features, resulting in instancevarying feature extraction. The output features of the search image can be directly used for predicting target locations without extra correlation step. Moreover, our model can be flexibly pre-trained on abundant unpaired images, leading to notably faster convergence than the existing methods. Extensive experiments show our method achieves the stateof-the-art results while running at real-time. Our feature networks also can be applied to existing tracking pipelines seamlessly to raise the tracking performance. 𝑓 𝑧 z (a1) Siamese-like feature network x Correlation Operation (b1) Feature correlation (a2) Target-dependent feature network 𝑓 𝑧 𝑓 𝑥 (b2) Ours z x 𝑓 cor 𝑓 𝑥 𝑓 cor (c) Prediction Localization Size estimation Prediction head 𝑓 cor
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu 等NeurIPS 2022 · 被引用 556 次
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo 等AAAI 2024 · 被引用 247 次
- MixFormerV2: Efficient Fully Transformer TrackingYutao Cui, Tianhui Song, Gangshan Wu, Limin WangNeurIPS 2023 · 被引用 193 次
- Robust Object Modeling for Visual TrackingYidong Cai, Jie Liu, Jie Tang, Gangshan WuICCV 2023 · 被引用 165 次
- Exploring Lightweight Hierarchical Vision Transformers for Efficient Visual TrackingBen Kang, Xin Chen, Dong Wang, Houwen Peng 等ICCV 2023 · 被引用 119 次
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang 等NeurIPS 2021 · 被引用 1,388 次
相关 Paper
- Reinforced Similarity Learning: Siamese Relation Networks for Robust Object TrackingDawei Zhang, Zhonglong Zheng, Minglu Li, Xiaowei He 等ACM MM 2020 · 被引用 15 次
- Deformable Siamese Attention Networks for Visual Object TrackingYuechen Yu, Yilei Xiong, Weilin Huang, Matthew R. ScottCVPR 2020
- Discriminative and Robust Online Learning for Siamese Visual TrackingJinghao Zhou, Peng Wang, Haoyang SunAAAI 2020 · 被引用 66 次
- Graph Attention TrackingDongyan Guo, Yanyan Shao, Ying Cui, Zhenhua Wang 等CVPR 2021
- Deep Meta Learning for Real-Time Target-Aware Visual TrackingJanghoon Choi, Junseok Kwon, Kyoung Mu LeeICCV 2019 · 被引用 93 次
