MixFormerV2: Efficient Fully Transformer Tracking
Yutao Cui, Tianhui Song, Gangshan Wu, Limin Wang
摘要
Transformer-based trackers have achieved strong accuracy on the standard benchmarks. However, their efficiency remains an obstacle to practical deployment on both GPU and CPU platforms. In this paper, to overcome this issue, we propose a fully transformer tracking framework, coined as MixFormerV2, without any dense convolutional operation and complex score prediction module. Our key design is to introduce four special prediction tokens and concatenate them with the tokens from target template and search areas. Then, we apply the unified transformer backbone on these mixed token sequence. These prediction tokens are able to capture the complex correlation between target template and search area via mixed attentions. Based on them, we can easily predict the tracking box and estimate its confidence score through simple MLP heads. To further improve the efficiency of MixFormerV2, we present a new distillation-based model reduction paradigm, including dense-to-sparse distillation and deep-to-shallow distillation. The former one aims to transfer knowledge from the dense-head based MixViT to our fully transformer tracker, while the latter one is used to prune some layers of the backbone. We instantiate two types of MixForemrV2, where the MixFormerV2-B achieves an AUC of 70.6% on LaSOT and an AUC of 57.4% on TNL2k with a high GPU speed of 165 FPS, and the MixFormerV2-S surpasses FEAR-L by 2.7% AUC on LaSOT with a real-time CPU speed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Cross-modulated Attention Transformer for RGBT TrackingYun Xiao, Jiacong Zhao, Andong Lu, Chenglong Li 等AAAI 2025 · 被引用 28 次
- Two-stream Beats One-stream: Asymmetric Siamese Network for Efficient Visual TrackingJiawen Zhu, Huayi Tang, Xin Chen, Xinying Wang 等AAAI 2025 · 被引用 27 次
- ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelYiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang 等NeurIPS 2024 · 被引用 21 次
- SUTrack: Towards Simple and Unified Single Object TrackingXin Chen, Ben Kang, Wanting Geng, Jiawen Zhu 等AAAI 2025 · 被引用 12 次
- General Compression Framework for Efficient Transformer Object TrackingLingyi Hong, Jinglun Li, Xinyu Zhou, Shilin Yan 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
相关 Paper
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 被引用 746 次
- FARTrack: Fast Autoregressive Visual Tracking with High PerformanceGuijie Wang, Tong Lin, Yifan Bai, Anjia Cao 等ICLR 2026 · 被引用 3 次
- Compact Transformer Tracker with Correlative Masked ModelingZikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen 等AAAI 2023 · 被引用 136 次
- Adaptive and Background-Aware Vision Transformer for Real-Time UAV TrackingShuiwang Li, Xiangxyang Yang, Dan Zeng, Xucheng WangICCV 2023 · 被引用 74 次
- Transforming Model Prediction for TrackingChristoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul 等CVPR 2022 · 被引用 399 次
