Joint 3D Object Detection and Tracking Using Spatio-Temporal Representation of Camera Image and LiDAR Point Clouds
Junho Koh, Jaekyum Kim, Jin Hyeok Yoo, Yecheol Kim, Dongsuk Kum, Jun Won Choi
Abstract
In this paper, we propose a new joint object detection and tracking (JoDT) framework for 3D object detection and tracking based on camera and LiDAR sensors. The proposed method, referred to as 3D DetecTrack, enables the detector and tracker to cooperate to generate a spatio-temporal representation of the camera and LiDAR data, with which 3D object detection and tracking are then performed. The detector constructs the spatio-temporal features via the weighted temporal aggregation of the spatial features obtained by the camera and LiDAR fusion. Then, the detector reconfigures the initial detection results using information from the tracklets maintained up to the previous time step. Based on the spatio-temporal features generated by the detector, the tracker associates the detected objects with previously tracked objects using a graph neural network (GNN). We devise a fully-connected GNN facilitated by a combination of rule-based edge pruning and attention-based edge gating, which exploits both spatial and temporal object contexts to improve tracking performance. The experiments conducted on both KITTI and nuScenes benchmarks demonstrate that the proposed 3D DetecTrack achieves significant improvements in both detection and tracking performances over baseline methods and achieves state-of-the-art performance among existing methods through collaboration between the detector and tracker.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0e1c48c-3059-4d60-82f4-68eedfe17062Cited by top-tier papers4
- Multi-Scene Generalized Trajectory Global Graph Solver with Composite Nodes for Multiple Object TrackingYan Gao, Haojun Xu, Jie Li, Nannan Wang et al.AAAI 2024 · 17 citations
- msLPCC: A Multimodal-Driven Scalable Framework for Deep LiDAR Point Cloud CompressionMiaohui Wang, Runnan Huang, Hengjin Dong, Di Lin et al.AAAI 2024 · 7 citations
- 3D Measurement of Complex Textured Objects Based on Bidirectional Fringe ProjectionYuchong Chen, Jian Yu, Shaoyan Gai, Zeyu Cai et al.AAAI 2025 · 4 citations
- High-Precision 3D Measurement of Complex Textured Surfaces Using Multiple Filtering ApproachYuchong Chen, Jian Yu, Shaoyan Gai, Zeyu Cai et al.ICCV 2025
Builds on7
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen et al.ICCV 2019 · 840 citations
- CIA-SSD: Confident IoU-Aware Single-Stage Object Detector From Point CloudWu Zheng, Weiliang Tang, Sijin Chen, Li Jiang et al.AAAI 2021 · 335 citations
- Joint Monocular 3D Vehicle Detection and TrackingHou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin et al.ICCV 2019 · 242 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
- RetinaTrack: Online Single Stage Joint Detection and TrackingZhichao Lu, Vivek Rathod, Ronny Votel, Jonathan HuangCVPR 2020
Related papers
- GNN3DMOT: Graph Neural Network for 3D Multi-Object Tracking With 2D-3D Multi-Feature LearningXinshuo Weng, Yongxin Wang, Yunze Man, Kris M. KitaniCVPR 2020
- PC-RGNN: Point Cloud Completion and Graph Neural Network for 3D Object DetectionYanan Zhang, Di Huang, Yunhong WangAAAI 2021 · 109 citations
- Point-GNN: Graph Neural Network for 3D Object Detection in a Point CloudWeijing Shi, Raj RajkumarCVPR 2020
- Exploring Simple 3D Multi-Object Tracking for Autonomous DrivingChenxu Luo, Xiaodong Yang, Alan L. YuilleICCV 2021 · 122 citations
- GraphAlign: Enhancing Accurate Feature Alignment by Graph matching for Multi-Modal 3D Object DetectionZiying Song, Haiyue Wei, Lin Bai, Lei Yang et al.ICCV 2023 · 73 citations
