Object-Centric Multiple Object Tracking
Zixu Zhao, Jiaze Wang, Max Horn, Yizhuo Ding, Tong He, Zechen Bai, Dominik Zietlow, Carl-Johann Simon-Gabriel, Bing Shuai, Zhuowen Tu, Thomas Brox, Bernt Schiele
Abstract
Unsupervised object-centric learning methods allow the partitioning of scenes into entities without additional localization information and are excellent candidates for reducing the annotation burden of multiple-object tracking (MOT) pipelines. Unfortunately, they lack two key properties: objects are often split into parts and are not consistently tracked over time. In fact, state-of-the-art models achieve pixel-level accuracy and temporal consistency by relying on supervised object detection with additional ID labels for the association through time. This paper proposes a video object-centric model for MOT. It consists of an index-merge module that adapts the object-centric slots into detection outputs and an object memory module that builds complete object prototypes to handle occlusions. Benefited from object-centric learning, we only require sparse detection labels (0%-6.25%) for object localization and feature binding. Relying on our self-supervised Expectation-Maximization-inspired loss for object association, our approach requires no ID labels. Our experiments significantly narrow the gap between the existing object-centric model and the fully supervised state-of-the-art and outperform several unsupervised trackers. Code is available at https://github.com/amazon-science/object-centric-multiple-object-tracking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8eb63ece-3a4b-41ca-b294-cb208b9b76f0Cited by top-tier papers4
- Ock: Unsupervised Dynamic Video Prediction With Object-Centric KinematicsYeon-Ji Song, Jaein Kim, Suhyung Choi, Jin-Hwa Kim et al.ICCV 2025 · 4 citations
- InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-LocalizationHongyang ZHANG, Maonan Wang, Ziyao Wang, Hongrui Yin et al.ICML 2026 · 1 citation
- LOMM: Latest Object Memory Management for Temporally Consistent Video Instance SegmentationSeunghun Lee, Jiwan Seo, Minwoo Choi, Kiljoon Han et al.ICCV 2025 · 1 citation
- Impossible VideosZechen Bai, Hai Ci, Mike Zheng ShouICML 2025
Builds on30
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani et al.ICLR 2021 · 357 citations
Related papers
- Self-supervised Object-Centric Learning for VideosGörkay Aydemir, Weidi Xie, Fatma GüneyNeurIPS 2023 · 61 citations
- VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object TrackingZekun Qian, Ruize Han, Junhui Hou, Linqi Song et al.ICCV 2025 · 3 citations
- Tracking without Label: Unsupervised Multiple Object Tracking via Contrastive Similarity LearningSha Meng, Dian Shao, Jiacheng Guo, Shan GaoICCV 2023 · 14 citations
- Self-Supervised Multi-Object Tracking with Cross-input ConsistencyFavyen Bastani, Songtao He, Samuel MaddenNeurIPS 2021 · 39 citations
- Temporally Consistent Object-Centric Learning by Contrasting SlotsAnna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius et al.CVPR 2025
