Transformer Tracking with Cyclic Shifting Window Attention
Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang
摘要
Transformer architecture has been showing its great strength in visual object tracking, for its effective attention mechanism. Existing transformer-based approaches adopt the pixel-to-pixel attention strategy on flattened image features and unavoidably ignore the integrity of ob-jects. In this paper, we propose a new transformer ar-chitecture with multi-scale cyclic shifting window attention for visual object tracking, elevating the attention from pixel to window level. The cross-window multi-scale at-tention has the advantage of aggregating attention at dif-ferent scales and generates the best fine-scale match for the target object. Furthermore, the cyclic shifting strat-egy brings greater accuracy by expanding the window sam-ples with positional information, and at the same time saves huge amounts of computational power by removing redun-dant calculations. Extensive experiments demonstrate the superior performance of our method, which also sets the new state-of-the-art records on five challenging datasets, along with the VOT2020, UAV123, LaSOT, TrackingNet, and GOT-lOk benchmarks. Our project is available at https://github.com/SkyeSong38/CSWinTT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper44
- MixFormerV2: Efficient Fully Transformer TrackingYutao Cui, Tianhui Song, Gangshan Wu, Limin WangNeurIPS 2023 · 被引用 193 次
- Compact Transformer Tracker with Correlative Masked ModelingZikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen 等AAAI 2023 · 被引用 136 次
- Exploring Lightweight Hierarchical Vision Transformers for Efficient Visual TrackingBen Kang, Xin Chen, Dong Wang, Houwen Peng 等ICCV 2023 · 被引用 119 次
- Adaptive and Background-Aware Vision Transformer for Real-Time UAV TrackingShuiwang Li, Xiangxyang Yang, Dan Zeng, Xucheng WangICCV 2023 · 被引用 74 次
- Foreground-Background Distribution Modeling Transformer for Visual Object TrackingDawei Yang, Jianfeng He, Yinchao Ma, Qianjin Yu 等ICCV 2023 · 被引用 44 次
它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan 等AAAI 2020 · 被引用 944 次
相关 Paper
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 被引用 746 次
- Target-Aware Tracking with Long-Term Context AttentionKaijie He, Canlong Zhang, Sheng Xie, Zhixin Li 等AAAI 2023 · 被引用 102 次
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 被引用 180 次
- Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual TrackingNing Wang, Wengang Zhou, Jie Wang, Houqiang LiCVPR 2021
- High-Performance Discriminative Tracking with TransformersBin Yu, Ming Tang, Linyu Zheng, Guibo Zhu 等ICCV 2021 · 被引用 113 次
