LA-MOTR: End-to-End Multi-Object Tracking by Learnable Association
Peng Wang, Yongcai Wang, Hualong Cao, Wang Chen, Deying Li
Abstract
This paper proposes LA-MOTR, a novel Tracking-by-Learnable-Association framework that resolves the competing optimization objectives between detection and association in end-to-end Tracking-by-Attention (TbA) Multi-Object Tracking. Current TbA methods employ shared decoders for simultaneous object detection and tracklet association, often resulting in task interference and suboptimal accuracy. By contrast, our end-to-end framework decouples these tasks into two specialized modules: Separated Object-Tracklet Detection (SOTD) and Spatial-Guided Learnable Association (SGLA). This decoupled design offers flexibility and explainability. In particular, SOTD independently detects new objects and existing tracklets in each frame, while SGLA associates them via Spatial-Weighted Learnable Attention module guided by relative spatial cues. Temporal coherence is further maintained through Tracklet Updates Module. The learnable association mechanism resolves the inherent suboptimal association issues in decoupled frameworks, avoiding the task interference commonly observed in joint approaches. Evaluations on DanceTrack, MOT17, and SportMOT datasets demonstrate state-of-theart performance. Extensive ablation studies validate the effectiveness of the designed modules. Code is available at https://github.com/PenK1nG/LA-MOTR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f653187c-40bd-4fdd-881d-b74ee95787d6Cited by top-tier papers2
- SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment AnythingPeng Wang, Yongcai Wang, Wang Chen, Hualong Cao et al.CVPR 2026
- Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World ScenesQi Zhang, Jixuan Chen, Zhang Kaiyi, Xinquan Yu et al.CVPR 2026
Builds on18
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
Related papers
- ADA-Track: End-to-End Multi-Camera 3D Multi-Object Tracking with Alternating Detection and AssociationShuxiao Ding, Lukas Schneider, Marius Cordts, Juergen GallCVPR 2024
- From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object TrackingYuqing Shao, Yuchen Yang, Rui Yu, Weilong Li et al.CVPR 2026 · 5 citations
- MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object TrackingRuopeng Gao, Limin WangICCV 2023 · 143 citations
- End-to-End Multiple Object Tracking with Dynamic Scene PerceptionRuonan Wei, Yuntao Wang, Siyan Fang, Yuehuan WangACM MM 2025 · 1 citation
- Learnable Graph Matching: Incorporating Graph Partitioning With Deep Feature Learning for Multiple Object TrackingJiawei He, Zehao Huang, Naiyan Wang, Zhaoxiang ZhangCVPR 2021
