Joint Spatial-Temporal Optimization for Stereo 3D Object Tracking
Peiliang Li, Jieqi Shi, Shaojie Shen
Abstract
Directly learning multiple 3D objects motion from sequential images is difficult, while the geometric bundle adjustment lacks the ability to localize the invisible object centroid. To benefit from both the powerful object understanding skill from deep neural network meanwhile tackle precise geometry modeling for consistent trajectory estimation, we propose a joint spatial-temporal optimization-based stereo 3D object tracking method. From the network, we detect corresponding 2D bounding boxes on adjacent images and regress an initial 3D bounding box. Dense object cues (local depth and local coordinates) that associating to the object centroid are then predicted using a region-based network. Considering both the instant localization accuracy and motion consistency, our optimization models the relations between the object centroid and observed cues into a joint spatial-temporal error function. All historic cues will be summarized to contribute to the current estimation by a per-frame marginalization strategy without repeated computation. Quantitative evaluation on the KITTI tracking dataset shows our approach outperforms previous imagebased 3D tracking methods by significant margins. We also report extensive results on multiple categories and larger datasets (KITTI raw and Argoverse Tracking) for future benchmarking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- You Don't Only Look Once: Constructing Spatial-Temporal Memory for Integrated 3D Object Detection and TrackingJiaming Sun, Yiming Xie, Siyu Zhang, Linghao Chen et al.ICCV 2021 · 12 citations
- Offboard 3D Object Detection From Point Cloud SequencesCharles R. Qi, Yin Zhou, Mahyar Najibi, Pei Sun et al.CVPR 2021
Builds on5
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous DrivingXinzhu Ma, Zhihui Wang, Haojie Li, Pengbo Zhang et al.ICCV 2019 · 339 citations
- Joint Monocular 3D Vehicle Detection and TrackingHou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin et al.ICCV 2019 · 242 citations
- FAMNet: Joint Learning of Feature, Affinity and Multi-Dimensional Assignment for Online Multiple Object TrackingPeng Chu, Haibin LingICCV 2019 · 229 citations
- Robust Multi-Modality Multi-Object TrackingWenwei Zhang, Hui Zhou, Shuyang Sun, Zhe Wang et al.ICCV 2019 · 221 citations
Related papers
- Stereo Neural Vernier CaliperShichao Li, Zechun Liu, Zhiqiang Shen, Kwang-Ting ChengAAAI 2022 · 6 citations
- Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual OdometryZhaoxing Zhang, Junda Cheng, Gangwei Xu, Xiaoxiang Wang et al.AAAI 2025 · 9 citations
- DSGN: Deep Stereo Geometry Network for 3D Object DetectionYilun Chen, Shu Liu, Xiaoyong Shen, Jiaya JiaCVPR 2020
- LIGA-Stereo: Learning LiDAR Geometry Aware Representations for Stereo-based 3D DetectorXiaoyang Guo, Shaoshuai Shi, Xiaogang Wang, Hongsheng LiICCV 2021 · 132 citations
- Feature-Level Collaboration: Joint Unsupervised Learning of Optical Flow, Stereo Depth and Camera MotionCheng Chi, Qingjie Wang, Tianyu Hao, Peng Guo et al.CVPR 2021
