Time3D: End-to-End Joint Monocular 3D Object Detection and Tracking for Autonomous Driving
Peixuan Li, Jieyu Jin
Abstract
While separately leveraging monocular 3D object detection and 2D multi-object tracking can be straightforwardly applied to sequence images in a frame-by-frame fashion, stand-alone tracker cuts off the transmission of the uncertainty from the 3D detector to tracking while cannot pass tracking error differentials back to the 3D detector. In this work, we propose jointly training 3D detection and 3D tracking from only monocular videos in an end-to-end manner. The key component is a novel spatial-temporal information flow module that aggregates geometric and appearance features to predict robust similarity scores across all objects in current and past frames. Specifically, we leverage the attention mechanism of the transformer, in which self-attention aggregates the spatial information in a specific frame, and cross-attention exploits relation and affinities of all objects in the temporal domain of sequence frames. The affinities are then supervised to estimate the trajectory and guide the flow of information between corresponding 3D objects. In addition, we propose a temporal -consistency loss that explicitly involves 3D target motion modeling into the learning, making the 3D trajectory smooth in the world coordinate system. Time3D achieves 21.4% AMOTA, 13.6% AMOTP on the nuScenes 3D tracking benchmark, surpassing all published competitors, and running at 38 FPS, while Time3D achieves 31.2% mAP, 39.4% NDS on the nuScenes 3D detection benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86c0b14b-4dd2-4ea0-90fb-eedc5567b704Cited by top-tier papers11
- End-to-end 3D Tracking with Decoupled QueriesYanwei Li, Zhiding Yu, Jonah Philion, Anima Anandkumar et al.ICCV 2023 · 32 citations
- Delving into Motion-Aware Matching for Monocular 3D Object TrackingKuan-Chih Huang, Ming-Hsuan Yang, Yi-Hsuan TsaiICCV 2023 · 20 citations
- Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed TransformerWenhan Wu, Ce Zheng, Zihao Yang, Chen Chen et al.ACM MM 2024 · 16 citations
- Predict to Detect: Prediction-guided 3D Object Detection using Sequential ImagesSanmin Kim, Youngseok Kim, In-Jae Lee, Dongsuk KumICCV 2023 · 16 citations
- Look More but Care Less in Video RecognitionYitian Zhang, Yue Bai, Huan Wang, Yi Xu et al.NeurIPS 2022 · 13 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Spatial-Temporal Relation Networks for Multi-Object TrackingJiarui Xu, Yue Cao, Zheng Zhang, Han HuICCV 2019 · 260 citations
- Joint Monocular 3D Vehicle Detection and TrackingHou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin et al.ICCV 2019 · 242 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
Related papers
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li et al.NeurIPS 2022 · 401 citations
- Rethinking the Spatio-Temporal Alignment of End-to-End 3D PerceptionXiaoyu Li, Peidong Li, Xian Wu, Long Shi et al.AAAI 2026 · 1 citation
- Cross Modal Transformer: Towards Fast and Robust 3D Object DetectionJunjie Yan, Yingfei Liu, Jianjian Sun, Fan Jia et al.ICCV 2023 · 143 citations
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
- 3DPPE: 3D Point Positional Encoding for Transformer-based Multi-Camera 3D Object DetectionChangyong Shu, Jiajun Deng, Fisher Yu, Yifan LiuICCV 2023 · 34 citations
