Multiple Object Tracking With Correlation Learning
Qiang Wang, Yun Zheng, Pan Pan, Yinghui Xu
Abstract
Recent works have shown that convolutional networks have substantially improved the performance of multiple object tracking by simultaneously learning detection and appearance features. However, due to the local perception of the convolutional network structure itself, the long-range dependencies in both the spatial and temporal cannot be obtained efficiently. To incorporate the spatial layout, we propose to exploit the local correlation module to model the topological relationship between targets and their surrounding environment, which can enhance the discriminative power of our model in crowded scenes. Specifically, we establish dense correspondences of each spatial location and its context, and explicitly constrain the correlation volumes through self-supervised learning. To exploit the temporal context, existing approaches generally utilize two or more adjacent frames to construct an enhanced feature representation, but the dynamic motion scene is inherently difficult to depict via CNNs. Instead, our paper proposes a learnable correlation operator to establish frameto-frame matches over convolutional feature maps in the different layers to align and propagate temporal context. With extensive experimental results on the MOT datasets, our approach demonstrates the effectiveness of correlation learning with the superior performance and obtains stateof-the-art MOTA of 76.5% and IDF1 of 73.6% on MOT17.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af47afcd-d5fa-4483-94f5-ba4fa2c0b548Cited by top-tier papers23
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- MeMOT: Multi-Object Tracking with MemoryJiarui Cai, Mingze Xu, Wei Li, Yuanjun Xiong et al.CVPR 2022 · 216 citations
- SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports ScenesYutao Cui, Chenkai Zeng, Xiaoyu Zhao, Yichun Yang et al.ICCV 2023 · 187 citations
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 180 citations
- MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?Matteo Fabbri, Guillem Brasó, Gianluca Maugeri, Orcun Cetintas et al.ICCV 2021 · 128 citations
Builds on10
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Spatial-Temporal Relation Networks for Multi-Object TrackingJiarui Xu, Yue Cao, Zheng Zhang, Han HuICCV 2019 · 260 citations
- FAMNet: Joint Learning of Feature, Affinity and Multi-Dimensional Assignment for Online Multiple Object TrackingPeng Chu, Haibin LingICCV 2019 · 229 citations
- Lifted Disjoint Paths with Application in Multiple Object TrackingAndrea Hornáková, Roberto Henschel, Bodo Rosenhahn, Paul SwobodaICML 2020 · 131 citations
- RetinaTrack: Online Single Stage Joint Detection and TrackingZhichao Lu, Vivek Rathod, Ronny Votel, Jonathan HuangCVPR 2020
Related papers
- MotionTrack: Learning Robust Short-Term and Long-Term Motions for Multi-Object TrackingZheng Qin, Sanping Zhou, Le Wang, Jinghai Duan et al.CVPR 2023
- Correlation-Aware Deep TrackingFei Xie, Chunyu Wang, Guangting Wang, Yue Cao et al.CVPR 2022 · 189 citations
- Exploiting Better Feature Aggregation for Video Object DetectionLiang Han, Pichao Wang, Zhaozheng Yin, Fan Wang et al.ACM MM 2020 · 37 citations
- LA-MOTR: End-to-End Multi-Object Tracking by Learnable AssociationPeng Wang, Yongcai Wang, Hualong Cao, Wang Chen et al.ICCV 2025 · 9 citations
- End-to-End Multiple Object Tracking with Dynamic Scene PerceptionRuonan Wei, Yuntao Wang, Siyan Fang, Yuehuan WangACM MM 2025 · 1 citation
