SUTrack: Towards Simple and Unified Single Object Tracking
Xin Chen, Ben Kang, Wanting Geng, Jiawen Zhu, Yi Liu, Dong Wang, Huchuan Lu
Abstract
In this paper, we propose a simple yet unified single object tracking (SOT) framework, dubbed SUTrack. It consolidates five SOT tasks (RGB-based, RGB-Depth, RGB-Thermal, RGB-Event, RGB-Language Tracking) into a unified model trained in a single session. Due to the distinct nature of the data, current methods typically design individual architectures and train separate models for each task. This fragmentation results in redundant training processes, repetitive technological innovations, and limited cross-modal knowledge sharing. In contrast, SUTrack demonstrates that a single model with a unified input representation can effectively handle various common SOT tasks, eliminating the need for task-specific designs and separate training sessions. Additionally, we introduce a task-recognition auxiliary training strategy and a soft token type embedding to further enhance SUTrack's performance with minimal overhead. Experiments show that SUTrack outperforms previous task-specific counterparts across 11 datasets spanning five SOT tasks. Moreover, we provide a range of models catering edge devices as well as high-performance GPUs, striking a good trade-off between speed and accuracy. We hope SUTrack could serve as a strong foundation for further compelling research into unified tracking models. Code and models are available at github.com/chenxin-dlut/SUTrack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d637018-617c-41c0-956f-4f4fc9d71b65Cited by top-tier papers13
- UTPTrack: Towards Simple and Unified Token Pruning for Visual TrackingHao Wu, Xudong Wang, Jialiang Zhang, Junlong Tong et al.CVPR 2026 · 6 citations
- General Compression Framework for Efficient Transformer Object TrackingLingyi Hong, Jinglun Li, Xinyu Zhou, Shilin Yan et al.ICCV 2025 · 5 citations
- What You Have is What You Track: Adaptive and Robust Multimodal TrackingYuedong Tan, Jiawei Shao, Eduard Zamfir, Ruanjun Li et al.ICCV 2025 · 5 citations
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingXiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang et al.ICCV 2025 · 3 citations
- Adaptive Depth Lightweight RGB-T Tracking with Holistic Token RoutingTian Ding, Hongtao Yang, Liangtao Shi, Jun Li et al.CVPR 2026 · 3 citations
Builds on50
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
Related papers
- UETrack: A Unified and Efficient Framework for Single Object TrackingBen Kang, Jie Zhao, Xin Chen, Wanting Geng et al.CVPR 2026 · 2 citations
- Single-Model and Any-Modality for Video Object TrackingZongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu et al.CVPR 2024 · 78 citations
- OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient TuningLingyi Hong, Shilin Yan, Renrui Zhang, Wanyun Li et al.CVPR 2024
- Unified Transformer Tracker for Object TrackingFan Ma, Mike Zheng Shou, Linchao Zhu, Haoqi Fan et al.CVPR 2022 · 121 citations
- Do Different Tracking Tasks Require Different Appearance Models?Zhongdao Wang, Hengshuang Zhao, Ya-Li Li, Shengjin Wang et al.NeurIPS 2021 · 107 citations
