Parametric Depth Based Feature Representation Learning for Object Detection and Segmentation in Bird's-Eye View
Jiayu Yang, Enze Xie, Miaomiao Liu, José M. Álvarez
Abstract
Recent vision-only perception models for autonomous driving achieved promising results by encoding multi-view image features into Bird’s-Eye-View (BEV) space. A critical step and the main bottleneck of these methods is transforming image features into the BEV coordinate frame. This paper focuses on leveraging geometry information, such as depth, to model such feature transformation. Existing works rely on non-parametric depth distribution modeling leading to significant memory consumption, or ignore the geometry information to address this problem. In contrast, we propose to use parametric depth distribution modeling for feature transformation. We first lift the 2D image features to the 3D space defined for the ego vehicle via a predicted parametric depth distribution for each pixel in each view. Then, we aggregate the 3D feature volume based on the 3D space occupancy derived from depth to the BEV frame. Finally, we use the transformed features for downstream tasks such as object detection and semantic segmentation. Existing semantic segmentation methods do also suffer from an hallucination problem as they do not take visibility information into account. This hallucination can be particularly problematic for subsequent modules such as control and planning. To mitigate the issue, our method provides depth uncertainty and reliable visibility-aware estimations. We further leverage our parametric depth modeling to present a novel visibility-aware evaluation metric that, when taken into account, can mitigate the hallucination problem. Ex tensive experiments on object detection and semantic segmentation on the nuScenes datasets demonstrate that our method outperforms existing methods on both tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 855c1478-4ddb-4e9f-b924-6586ef31ce19Cited by top-tier papers5
- SG-BEV: Satellite-Guided BEV Fusion for Cross-View Semantic SegmentationJunyan Ye, Qiyan Luo, Jinhua Yu, Huaping Zhong et al.CVPR 2024 · 19 citations
- BevSplat: Resolving Height Ambiguity via Feature-Based Gaussian Primitives for Weakly-Supervised Cross-View LocalizationQiwei Wang, Shaoxun Wu, Yujiao ShiNeurIPS 2025 · 10 citations
- PointBeV: A Sparse Approach to BeV PredictionsLoïck Chambon, Éloi Zablocki, Mickaël Chen, Florent Bartoccioni et al.CVPR 2024
- Leveraging 3D Geometric Priors in 2D Rotation Symmetry DetectionAhyun Seo, Minsu ChoCVPR 2025
- SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large ObjectsAbhinav Kumar, Yuliang Guo, Xinyu Huang, Liu Ren et al.CVPR 2024
Builds on12
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera et al.ICCV 2019 · 504 citations
- Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object DetectionJinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer et al.ICLR 2023 · 71 citations
Related papers
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
- SA-BEV: Generating Semantic-Aware Bird's-Eye-View Feature for Multi-view 3D Object DetectionJinqing Zhang, Yanan Zhang, Qingjie Liu, Yunhong WangICCV 2023 · 41 citations
- GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy PredictionYuanhui Huang, Amonnut Thammatadatrakoon, Wenzhao Zheng, Yunpeng Zhang et al.CVPR 2025
- Tri-Perspective View for Vision-Based 3D Semantic Occupancy PredictionYuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou et al.CVPR 2023
- CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird’s-Eye-View Semantic SegmentationJeongbin Hong, Dooseop Choi, Taeg-Hyun An, KYOUNG AN AN et al.CVPR 2026 · 1 citation
