Delving into Motion-Aware Matching for Monocular 3D Object Tracking
Kuan-Chih Huang, Ming-Hsuan Yang, Yi-Hsuan Tsai
Abstract
Recent advances of monocular 3D object detection facilitate the 3D multi-object tracking task based on lowcost camera sensors. In this paper, we find that the motion cue of objects along different time frames is critical in 3D multi-object tracking, which is less explored in existing monocular-based approaches. To this end, we propose MoMA-M3T, a framework that mainly consists of three motion-aware components. First, we represent the possible movement of an object related to all object tracklets in the feature space as its motion features. Then, we further model the historical object tracklet along the time frame in a spatial-temporal perspective via a motion transformer. Finally, we propose a motion-aware matching module to associate historical object tracklets and current observations as final tracking results. We conduct extensive experiments on the nuScenes and KITTI datasets to demonstrate that our MoMA-M3T achieves competitive performance against state-of-the-art methods. Moreover, the proposed tracker is flexible and can be easily plugged into existing imagebased 3D object detectors without re-training. Code and models are available at https:// github.com/ kuanchihhuang/ MoMA-M3T.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a952746-9c58-4f91-9817-b1b7c2d5e205Cited by top-tier papers5
- VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object TrackingZekun Qian, Ruize Han, Junhui Hou, Linqi Song et al.ICCV 2025 · 3 citations
- Category-Specific Selective Feature Enhancement for Long-Tailed Multi-Label Image ClassificationRuiqi Du, Xu Tang, Xiangrong Zhang, Jingjing MaICCV 2025 · 1 citation
- COVTrack: Continuous Open-Vocabulary Tracking via Adaptive Multi-Cue FusionZekun Qian, Ruize Han, Zhixiang Wang, Junhui Hou et al.ICCV 2025 · 1 citation
- Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object TrackingHaonan Zhang, Xinyao Wang, Boxi Wu, Tu Zheng et al.AAAI 2026
- Instantaneous Perception of Moving Objects in 3DDi Liu, Bingbing Zhuang, Dimitris N. Metaxas, Manmohan ChandrakerCVPR 2024
Builds on25
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera et al.ICCV 2019 · 504 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
Related papers
- Time3D: End-to-End Joint Monocular 3D Object Detection and Tracking for Autonomous DrivingPeixuan Li, Jieyu JinCVPR 2022 · 52 citations
- Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object TrackingZiqi Pang, Jie Li, Pavel Tokmakov, Dian Chen et al.CVPR 2023
- TrajectoryFormer: 3D Object Tracking Transformer with Predictive Trajectory HypothesesXuesong Chen, Shaoshuai Shi, Chao Zhang, Benjin Zhu et al.ICCV 2023 · 25 citations
- GRAE-3DMOT: Geometry Relation-Aware Encoder for Online 3D Multi-Object TrackingHyunseop Kim, Hyo-Jun Lee, Yonguk Lee, Jinu Lee et al.CVPR 2025
- MeMOT: Multi-Object Tracking with MemoryJiarui Cai, Mingze Xu, Wei Li, Yuanjun Xiong et al.CVPR 2022 · 216 citations
