CO-MOT: Boosting End-to-end Transformer-based Multi-Object Tracking via Coopetition Label Assignment and Shadow Sets
Feng Yan, Weixin Luo, Yujie Zhong, Yiyang Gan, Lin Ma
Abstract
Existing end-to-end Multi-Object Tracking (e2e-MOT) methods have not surpassed non-end-to-end tracking-by-detection methods. One possible reason lies in the training label assignment strategy that consistently binds the tracked objects with tracking queries and assigns few newborns to detection queries. Such an assignment, with one-to-one bipartite matching, yields an unbalanced training, i.e., scarce positive samples for detection queries, especially for an enclosed scenew with the majority of the newborns at the beginning of videos. As such, e2e-MOT will incline to generate a tracking terminal without renewal or re-initialization, compared to other tracking-by-detection methods. To alleviate this problem, we propose Co-MOT, a simple yet effective method to facilitate e2e-MOT by a novel coopetition label assignment with a shadow concept. Specifically, we add tracked objects to the matching targets for detection queries when performing the label assignment for training the intermediate decoders. For query initialization, we expand each query by a set of shadow counterparts with limited disturbance to itself. With extensive ablation studies, Co-MOT achieves superior performances without extra costs, e.g., 69.4% HOTA on DanceTrack and 52.8% TETA on BDD100K. Impressively, Co-MOT only requires 38% FLOPs of MOTRv2 with comparable performances, resulting in the 1.4⇥ faster inference speed. Source code is publicly available at https://github.com/BingfengYan/CO-MOT . INTRODUCTION Multi-Object Tracking (MOT) is traditionally tackled by a series of tasks, e.g., object detection (Zhao
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02b740ed-775a-4d2a-893f-970d0489c610Cited by top-tier papers2
- Out of Sight, Out of Track: Adversarial Attacks on Propagation-based Multi-Object Trackers via Query State ManipulationHalima Bouzidi, Haoyu Liu, Yonatan Achamyeleh, Praneetsai Iddamsetty et al.CVPR 2026 · 1 citation
- GLoMOT: Efficient Online GNN-based Low-Frame-Rate Multi-Object TrackerYaxuan Hu, Jie Hua, Gang Wu, Yuhong Yang et al.AAAI 2026
Builds on24
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
- DETRs with Collaborative Hybrid Assignments TrainingZhuofan Zong, Guanglu Song, Yu LiuICCV 2023 · 594 citations
Related papers
- MOTRv2: Bootstrapping End-to-End Multi-Object Tracking by Pretrained Object DetectorsYuang Zhang, Tiancai Wang, Xiangyu ZhangCVPR 2023
- Collaborative Tracking Learning for Frame-Rate-Insensitive Multi-Object TrackingYiheng Liu, Junta Wu, Yi FuICCV 2023 · 17 citations
- LA-MOTR: End-to-End Multi-Object Tracking by Learnable AssociationPeng Wang, Yongcai Wang, Hualong Cao, Wang Chen et al.ICCV 2025 · 9 citations
- TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training ModelBo Pang, Yizhuo Li, Yifan Zhang, Muchen Li et al.CVPR 2020
- ADA-Track: End-to-End Multi-Camera 3D Multi-Object Tracking with Alternating Detection and AssociationShuxiao Ding, Lukas Schneider, Marius Cordts, Juergen GallCVPR 2024
