STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection
Huijie Fan, Pengrui Huang, Qiang Wang, Baojie Fan, Jiahua Dong, Liangqiong Qu
Abstract
Existing surrounding-view 3D object detectors initialize high-confidence queries using current 2D information, while leveraging historical 3D features as priors. However, such heavy reliance on 2D cues introduces spatio-temporal inconsistencies between 2D and 3D representations. Specifically, 2D cues lack sufficient spatial information, limiting 3D localization capability. Moreover, insufficient temporal interaction often leads to object omission under occlusion. To address these challenges, we propose STUR3D, a unified framework establishing spatio-temporal alignment between 2D and 3D perception. First, we project temporal 3D features to the 2D image plane, empowering the 2D detector to distill representations essential for 3D localization, harmonizing cross-dimensional information. Second, we inject temporal cues into 2D detection, fostering spatio-temporal reasoning, and ensuring robust 3D detection under dynamic scenes and occlusion. Additionally, we embed depth-aware geometric cues into features for 2D-to-3D lifting, mitigating inherent ambiguities. Extensive nuScenes experiments validate STUR3D, achieving SOTA on the test set with 57.9% mAP and 64.6% NDS. Our code will be released at https://github.com/snowindog/STUR3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b7a447f-6095-480d-a530-fa0da84186cbBuilds on14
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li et al.ICCV 2023 · 399 citations
- Far3D: Expanding the Horizon for Surround-View 3D Object DetectionXiaohui Jiang, Shuailin Li, Yingfei Liu, Shihao Wang et al.AAAI 2024 · 100 citations
- Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object DetectionJinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer et al.ICLR 2023 · 71 citations
Related papers
- Object as Query: Lifting any 2D Object Detector to 3D DetectionZitian Wang, Zehao Huang, Jiahui Fu, Naiyan Wang et al.ICCV 2023 · 47 citations
- Time3D: End-to-End Joint Monocular 3D Object Detection and Tracking for Autonomous DrivingPeixuan Li, Jieyu JinCVPR 2022 · 52 citations
- Viewpoint Equivariance for Multi-View 3D Object DetectionDian Chen, Jie Li, Vitor Guizilini, Rares Ambrus et al.CVPR 2023
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 · 175 citations
- Enhancing 3D Object Detection with 2D Detection-Guided Query AnchorsHaoxuanye Ji, Pengpeng Liang, Erkang ChengCVPR 2024
