Predict to Detect: Prediction-guided 3D Object Detection using Sequential Images
Sanmin Kim, Youngseok Kim, In-Jae Lee, Dongsuk Kum
Abstract
Recent camera-based 3D object detection methods have introduced sequential frames to improve the detection performance hoping that multiple frames would mitigate the large depth estimation error. Despite improved detection performance, prior works rely on naive fusion methods (e.g., concatenation) or are limited to static scenes (e.g., temporal stereo), neglecting the importance of the motion cue of objects. These approaches do not fully exploit the potential of sequential images and show limited performance improvements. To address this limitation, we propose a novel 3D object detection model, P2D (Predict to Detect), that integrates a prediction scheme into a detection framework to explicitly extract and leverage motion features. P2D predicts object information in the current frame using solely past frames to learn temporal motion features. We then introduce a novel temporal feature aggregation method that attentively exploits Bird’s-Eye-View (BEV) features based on predicted object information, resulting in accurate 3D object detection. Experimental results demonstrate that P2D improves mAP and NDS by 3.0% and 3.7% compared to the sequential image-based baseline, proving that incorporating a prediction scheme can significantly improve detection accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4ba40b8-48e4-4f78-804e-13a035001376Cited by top-tier papers4
- MemDistill: Distilling LiDAR Knowledge into Memory for Camera-Only 3D Object DetectionDonghyeon Kwon, Youngseok Yoon, Hyeongseok Son, Suha KwakICCV 2025 · 1 citation
- ForeSight: Multi-View Streaming Joint Object Detection and Trajectory ForecastingSandro Papais, Letian Wang, Brian Cheong, Steven L. WaslanderICCV 2025 · 1 citation
- MonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object DetectionHou-I Liu, Christine Wu, Jen-Hao Cheng, Wenhao Chai et al.CVPR 2025
- SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large ObjectsAbhinav Kumar, Yuliang Guo, Xinyu Huang, Liu Ren et al.CVPR 2024
Builds on20
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
Related papers
- Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object PredictionZhuofan Zong, Dongzhi Jiang, Guanglu Song, Zeyue Xue et al.ICCV 2023 · 63 citations
- BEVNeXt: Reviving Dense BEV Frameworks for 3D Object DetectionZhenxin Li, Shiyi Lan, José M. Álvarez, Zuxuan WuCVPR 2024
- CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object DetectionJisong Kim, Minjae Seong, Jun Won ChoiNeurIPS 2024 · 27 citations
- Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting LearningYang Jiao, Zequn Jie, Shaoxiang Chen, Lechao Cheng et al.AAAI 2024 · 13 citations
- MGTANet: Encoding Sequential LiDAR Points Using Long Short-Term Motion-Guided Temporal Attention for 3D Object DetectionJunho Koh, Junhyung Lee, Youngwoo Lee, Jaekyum Kim et al.AAAI 2023 · 34 citations
