Tracking without Label: Unsupervised Multiple Object Tracking via Contrastive Similarity Learning
Sha Meng, Dian Shao, Jiacheng Guo, Shan Gao
Abstract
Unsupervised learning is a challenging task due to the lack of labels. Multiple Object Tracking (MOT), which inevitably suffers from mutual object interference, occlusion, etc., is even more difficult without label supervision. In this paper, we explore the latent consistency of sample features across video frames and propose an Unsupervised Contrastive Similarity Learning method, named UCSL, including three contrast modules: self-contrast, cross-contrast, and ambiguity contrast. Specifically, i) self-contrast uses intra-frame direct and inter-frame indirect contrast to obtain discriminative representations by maximizing self-similarity. ii) Cross-contrast aligns cross- and continuous-frame matching results, mitigating the persistent negative effect caused by object occlusion. And iii) ambiguity contrast matches ambiguous objects with each other to further increase the certainty of subsequent object association through an implicit manner. On existing benchmarks, our method outperforms the existing unsupervised methods using only limited help from ReID head, and even provides higher accuracy than lots of fully supervised methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fabe386-9451-45ff-8d11-8dc8d027d46aCited by top-tier papers5
- Self-Supervised Multi-Object Tracking with Path ConsistencyZijia Lu, Bing Shuai, Yanbei Chen, Zhenlin Xu et al.CVPR 2024 · 13 citations
- Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object TrackingGuangyao Li, Siping Zhuang, Yajun Jian, Yan Yan et al.ICCV 2025 · 8 citations
- Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured VideosJianbo Ma, Hui Luo, Qi Chen, Yuankai Qi et al.AAAI 2026 · 2 citations
- ASCENT: Annotation-Free Self-Supervised Contrastive Embeddings for 3D Neuron Tracking in Fluorescence MicroscopyHaejun Han, Hang LuICCV 2025 · 1 citation
- Temporally Consistent Object-Centric Learning by Contrasting SlotsAnna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius et al.CVPR 2025
Builds on9
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 356 citations
- Unsupervised Learning of Accurate Siamese TrackingQiuhong Shen, Lei Qiao, Jinyang Guo, Peixia Li et al.CVPR 2022 · 73 citations
- Towards Discriminative Representation: Multi-view Trajectory Contrastive Learning for Online Multi-object TrackingEn Yu, Zhuoling Li, Shoudong HanCVPR 2022 · 54 citations
- Self-Supervised Multi-Object Tracking with Cross-input ConsistencyFavyen Bastani, Songtao He, Samuel MaddenNeurIPS 2021 · 39 citations
Related papers
- Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyHaiping Wu, Xiaolong WangICCV 2021 · 35 citations
- Object-Centric Multiple Object TrackingZixu Zhao, Jiaze Wang, Max Horn, Yizhuo Ding et al.ICCV 2023 · 10 citations
- Contrastive Transformation for Self-supervised Correspondence LearningNing Wang, Wengang Zhou, Houqiang LiAAAI 2021 · 38 citations
- Uncertainty-aware Unsupervised Multi-Object TrackingKai Liu, Sheng Jin, Zhihang Fu, Ze Chen et al.ICCV 2023 · 21 citations
- Learning to Track Instances without Video AnnotationsYang Fu, Sifei Liu, Umar Iqbal, Shalini De Mello et al.CVPR 2021
