MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space Model
Changcheng Xiao, Qiong Cao, Zhigang Luo, Long Lan
摘要
Tracking by detection has been the prevailing paradigm in the field of Multi-object Tracking (MOT). These methods typically rely on the Kalman Filter to estimate the future locations of objects, assuming linear object motion. However, they fall short when tracking objects exhibiting nonlinear and diverse motion in scenarios like dancing and sports. In addition, there has been limited focus on utilizing learning-based motion predictors in MOT. To address these challenges, we resort to exploring data-driven motion prediction methods. Inspired by the great expectation of state space models (SSMs), such as Mamba, in long-term sequence modeling with near-linear complexity, we introduce a Mamba-based motion model named Mamba moTion Predictor (MTP). MTP is designed to model the complex motion patterns of objects like dancers and athletes. Specifically, MTP takes the spatial-temporal location dynamics of objects as input, captures the motion pattern using a bi-Mamba encoding layer, and predicts the next motion. In real-world scenarios, objects may be missed due to occlusion or motion blur, leading to premature termination of their trajectories. To tackle this challenge, we further expand the application of MTP. We employ it in an autoregressive way to compensate for missing observations by utilizing its own predictions as inputs, thereby contributing to more consistent trajectories. Our proposed tracker, MambaTrack, demonstrates advanced performance on benchmarks such as Dancetrack and SportsMOT, which are characterized by complex motion and severe occlusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou 等CVPR 2026 · 被引用 12 次
- Hypergraph-State Collaborative Reasoning for Multi-Object TrackingZikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 等CVPR 2026 · 被引用 4 次
- Dual-level Adaptation for Multi-Object Tracking: Building Test-Time Calibration from Experience and IntuitionWen Guo, Pengfei Zhao, Zongmeng Wang, Yufan Hu 等CVPR 2026 · 被引用 4 次
- Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object TrackingChunjiang Li, Jianbo Ma, Li Shen, Yanru Chen 等CVPR 2026 · 被引用 1 次
- When Trackers Date Fish: A Benchmark and Framework for Underwater Multiple Fish TrackingWeiran Li, Yeqiang Liu, Qiannan Guo, Yijie Wei 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper20
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra 等NeurIPS 2020 · 被引用 1,100 次
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 被引用 1,030 次
- DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse MotionPeize Sun, Jinkun Cao, Yi Jiang, Zehuan Yuan 等CVPR 2022 · 被引用 305 次
相关 Paper
- Samba: Synchronized Set-of-Sequences Modeling for Multiple Object TrackingMattia Segù, Luigi Piccinelli, Siyuan Li, Yung-Hsu Yang 等ICLR 2025
- DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear PredictionWeiyi Lv, Yuhang Huang, Ning Zhang, Ruei-Sung Lin 等CVPR 2024 · 被引用 36 次
- PlugTrack: Multi-Perceptive Motion Analysis for Adaptive Fusion in Multi-Object TrackingSeungjae Kim, SeungJoon Lee, MyeongAh ChoAAAI 2026
- Realistic Full-Body Motion Generation from Sparse Tracking with State Space ModelKun Dong, Jian Xue, Zehai Niu, Xing Lan 等ACM MM 2024 · 被引用 7 次
- MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language TrackingXinqi Liu, Li Zhou, Zikun Zhou, Jianqiu Chen 等CVPR 2025
