Unified Transformer Tracker for Object Tracking
Fan Ma, Mike Zheng Shou, Linchao Zhu, Haoqi Fan, Yilei Xu, Yi Yang, Zhicheng Yan
Abstract
As an important area in computer vision, object tracking has formed two separate communities that respectively study Single Object Tracking (SOT) and Multiple Object Tracking (MOT). However, current methods in one tracking scenario are not easily adapted to the other due to the divergent training datasets and tracking objects of both tasks. Although UniTrack [45] demonstrates that a shared appearance model with multiple heads can be used to tackle individual tracking tasks, it fails to exploit the large-scale tracking datasets for training and performs poorly on the single object tracking. In this work, we present the Unified Transformer Tracker (UTT) to address tracking problems in different scenarios with one paradigm. A track transformer is developed in our UTT to track the target in both SOT and MOT where the correlation between the target feature and the tracking frame feature is exploited to localize the target. We demonstrate that both SOT and MOT tasks can be solved within this framework, and the model can be simultaneously end-to-end trained by alternatively optimizing the SOT and MOT objectives on the datasets of individual tasks. Extensive experiments are conducted on several benchmarks with a unified model trained on both SOT and MOT datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fd2299a-a835-4cc1-ad64-2db5d5d2cd22Cited by top-tier papers17
- Reading Relevant Feature from Global Representation Memory for Visual Object TrackingXinyu Zhou, Pinxue Guo, Lingyi Hong, Jinglun Li et al.NeurIPS 2023 · 33 citations
- DiffusionTrack: Point Set Diffusion Model for Visual Object TrackingFei Xie, Zhongdao Wang, Chao MaCVPR 2024 · 30 citations
- Single-Stage Visual Query Localization in Egocentric VideosHanwen Jiang, Santhosh Kumar Ramakrishnan, Kristen GraumanNeurIPS 2023 · 27 citations
- Integrating Boxes and Masks: A Multi-Object Framework for Unified Visual Tracking and SegmentationYuanyou Xu, Zongxin Yang, Yi YangICCV 2023 · 18 citations
- Self-Supervised Multi-Object Tracking with Path ConsistencyZijia Lu, Bing Shuai, Yanbei Chen, Zhenlin Xu et al.CVPR 2024 · 13 citations
Builds on17
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan et al.AAAI 2020 · 944 citations
Related papers
- Do Different Tracking Tasks Require Different Appearance Models?Zhongdao Wang, Hengshuang Zhao, Ya-Li Li, Shengjin Wang et al.NeurIPS 2021 · 107 citations
- SUTrack: Towards Simple and Unified Single Object TrackingXin Chen, Ben Kang, Wanting Geng, Jiawen Zhu et al.AAAI 2025 · 12 citations
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 354 citations
- Improving Multiple Object Tracking With Single Object TrackingLinyu Zheng, Ming Tang, Yingying Chen, Guibo Zhu et al.CVPR 2021
- Dual-Path Temporal Decoder for End-to-End Multi-Object TrackingHyunseop Kim, Juheon Jeong, Hanul Kim, Yeong Jun KohNeurIPS 2025 · 4 citations
