BEVStereo: Enhancing Depth Estimation in Multi-View 3D Object Detection with Temporal Stereo
Yinhao Li, Han Bao, Zheng Ge, Jinrong Yang, Jianjian Sun, Zeming Li
Abstract
Restricted by the ability of depth perception, all Multi-view 3D object detection methods fall into the bottleneck of depth accuracy. By constructing temporal stereo, depth estimation is quite reliable in indoor scenarios. However, there are two difficulties in directly integrating temporal stereo into outdoor multi-view 3D object detectors: 1) The construction of temporal stereos for all views results in high computing costs. 2) Unable to adapt to challenging outdoor scenarios. In this study, we propose an effective method for creating temporal stereo by dynamically determining the center and range of the temporal stereo. The most confident center is found using the EM algorithm. Numerous experiments on nuScenes have shown the BEVStereo's ability to deal with complex outdoor scenarios that other stereo-based methods are unable to handle. For the first time, a stereo-based approach shows superiority in scenarios like a static ego vehicle and moving objects. BEVStereo achieves the new state-of-the-art in the cameraonly track of nuScenes dataset while maintaining memory efficiency. Codes have been released 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc6db2be-a683-4614-b73e-ea253d161a94Cited by top-tier papers7
- M-BEV: Masked BEV Perception for Robust Autonomous DrivingSiran Chen, Yue Ma, Yu Qiao, Yali WangAAAI 2024 · 24 citations
- Panopticus: Omnidirectional 3D Object Detection on Resource-constrained Edge DevicesJeho Lee, Chanyoung Jung, Jiwon Kim, Hojung ChaMobiCom 2024 · 4 citations
- FastRSR: Efficient and Accurate Road Surface Reconstruction in Bird's Eye ViewYuting Zhao, Yuheng Ji, Xiaoshuai Hao, Shuxiao LiACM MM 2025 · 2 citations
- EchoDiffusion: Waveform Conditioned Diffusion Models for Echo-Based Depth EstimationWenjie Zhang, Jun Yin, Long Ma, Peng Yu et al.AAAI 2025 · 2 citations
- SPHERE: Semantic-PHysical Engaged REpresentation for 3D Semantic Scene CompletionZhiwen Yang, Yuxin PengACM MM 2025
Builds on15
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- How Do Neural Networks See Depth in Single Images?Tom van Dijk, Guido de CroonICCV 2019 · 210 citations
Related papers
- Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting LearningYang Jiao, Zequn Jie, Shaoxiang Chen, Lechao Cheng et al.AAAI 2024 · 13 citations
- Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object DetectionJinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer et al.ICLR 2023 · 71 citations
- UniFusion: Unified Multi-view Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-ViewZequn Qin, Jingyu Chen, Chao Chen, Xiaozhi Chen et al.ICCV 2023 · 38 citations
- Unleashing the Temporal Potential of Stereo Event Cameras for Continuous-Time 3D Object DetectionJae-Young Kang, Hoonhee Cho, Kuk-Jin YoonICCV 2025 · 4 citations
- Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object PredictionZhuofan Zong, Dongzhi Jiang, Guanglu Song, Zeyue Xue et al.ICCV 2023 · 63 citations
