TrackFormer: Multi-Object Tracking with Transformers
Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph Feichtenhofer
Abstract
The challenging task of multi-object tracking (MOT) requires simultaneous reasoning about track initialization, identity, and spatio-temporal trajectories. We formulate this task as a frame-to-frame set prediction problem and introduce TrackFormer, an end-to-end trainable MOT approach based on an encoder-decoder Transformer architecture. Our model achieves data association between frames via attention by evolving a set of track predictions through a video sequence. The Transformer decoder initializes new tracks from static object queries and autoregressively follows existing tracks in space and time with the conceptually new and identity preserving track queries. Both query types benefit from self-and encoder-decoder attention on global frame-level features, thereby omitting any additional graph optimization or modeling of motion and/or appearance. TrackFormer introduces a new tracking-by-attention paradigm and while simple in its design is able to achieve state-of-the-art performance on the task of multi-object tracking (MOT17) and segmentation (MOTS20). The code is available at https://github . com/timmeinhardt/trackformer
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers165
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath et al.ICLR 2026 · 1,103 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li et al.ICCV 2023 · 399 citations
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
Builds on14
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Learning to Track with Object PermanencePavel Tokmakov, Jie Li, Wolfram Burgard, Adrien GaidonICCV 2021 · 241 citations
- FAMNet: Joint Learning of Feature, Affinity and Multi-Dimensional Assignment for Online Multiple Object TrackingPeng Chu, Haibin LingICCV 2019 · 229 citations
- Lifted Disjoint Paths with Application in Multiple Object TrackingAndrea Hornáková, Roberto Henschel, Bodo Rosenhahn, Paul SwobodaICML 2020 · 131 citations
Related papers
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 180 citations
- TGFormer: Transformer with Track Query Group for Multi-Object TrackingRui Zeng, Yuanzhou Huang, Songwei PeiAAAI 2025 · 6 citations
- 3DMOTFormer: Graph Transformer for Online 3D Multi-Object TrackingShuxiao Ding, Eike Rehder, Lukas Schneider, Marius Cordts et al.ICCV 2023 · 36 citations
- Dual-Path Temporal Decoder for End-to-End Multi-Object TrackingHyunseop Kim, Juheon Jeong, Hanul Kim, Yeong Jun KohNeurIPS 2025 · 4 citations
- MeMOT: Multi-Object Tracking with MemoryJiarui Cai, Mingze Xu, Wei Li, Yuanjun Xiong et al.CVPR 2022 · 216 citations
