Assignment-Space-based Multi-Object Tracking and Segmentation
Anwesa Choudhuri, Girish Chowdhary, Alexander G. Schwing
Abstract
Multi-object tracking and segmentation (MOTS) is important for understanding dynamic scenes in video data. Existing methods perform well on multi-object detection and segmentation for independent video frames, but tracking of objects over time remains a challenge. MOTS methods formulate tracking locally, i.e., frame-by-frame, leading to sub-optimal results. Classical global methods on tracking operate directly on object detections, which leads to a combinatorial growth in the detection space. In contrast, we formulate a global method for MOTS over the space of assignments rather than detections: First, we find all top-k assignments of objects detected and segmented between any two consecutive frames and develop a structured prediction formulation to score assignment sequences across any number of consecutive frames. We use dynamic programming to find the global optimizer of this formulation in polynomial time. Second, we connect objects which reappear after having been out of view for some time. For this we formulate an assignment problem. On the challenging KITTI-MOTS and MOTSChallenge datasets, this achieves state-of-the-art results among methods which don’t use depth data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45d631c5-61c5-49d8-97be-edfd9c356d0bCited by top-tier papers3
- Tracking Anything with Decoupled Video SegmentationHo Kei Cheng, Seoung Wug Oh, Brian L. Price, Alexander G. Schwing et al.ICCV 2023 · 240 citations
- OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and CaptioningAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingNeurIPS 2024 · 7 citations
- Context-Aware Relative Object Queries to Unify Video Instance and Panoptic SegmentationAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingCVPR 2023
Builds on4
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Lifted Disjoint Paths with Application in Multiple Object TrackingAndrea Hornáková, Roberto Henschel, Bodo Rosenhahn, Paul SwobodaICML 2020 · 131 citations
- VIP-DeepLab: Learning Visual Perception With Depth-Aware Video Panoptic SegmentationSiyuan Qiao, Yukun Zhu, Hartwig Adam, Alan L. Yuille et al.CVPR 2021
- Learning a Neural Solver for Multiple Object TrackingGuillem Brasó, Laura Leal-TaixéCVPR 2020
Related papers
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 180 citations
- Joint Spatial-Temporal Optimization for Stereo 3D Object TrackingPeiliang Li, Jieqi Shi, Shaojie ShenCVPR 2020
- Integrating Boxes and Masks: A Multi-Object Framework for Unified Visual Tracking and SegmentationYuanyou Xu, Zongxin Yang, Yi YangICCV 2023 · 18 citations
- Delving into Motion-Aware Matching for Monocular 3D Object TrackingKuan-Chih Huang, Ming-Hsuan Yang, Yi-Hsuan TsaiICCV 2023 · 20 citations
- Multiple Object Tracking as ID PredictionRuopeng Gao, Ji Qi, Limin WangCVPR 2025
