GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object Detection
Jinqing Zhang, Yanan Zhang, Yunlong Qi, Zehua Fu, Qingjie Liu, Yunhong Wang
Abstract
Bird's-Eye-View (BEV) representation has emerged as a mainstream paradigm for multi-view 3D object detection, demonstrating impressive perceptual capabilities. However, existing methods overlook the geometric quality of BEV representation, leaving it in a low-resolution state and failing to restore the authentic geometric information of the scene. In this paper, we identify the drawbacks of previous approaches that limit the geometric quality of BEV representation and propose Radial-Cartesian BEV Sampling (RC-Sampling), which outperforms other feature transformation methods in efficiently generating high-resolution dense BEV representation to restore fine-grained geometric information. Additionally, we design a novel In-Box Label to substitute the traditional depth label generated from the Li-DAR points. This label reflects the actual geometric structure of objects rather than just their surfaces, injecting realworld geometric information into the BEV representation. In conjunction with the In-Box Label, Centroid-Aware Inner Loss (CAI Loss) is developed to capture the inner geometric structure of objects. Finally, we integrate the aforementioned modules into a novel multi-view 3D object detector, dubbed GeoBEV, which achieves a state-of-the-art result of 66.2% NDS on the nuScenes test set. The code is available at https://github.com/mengtan00/GeoBEV.git . * Corresponding author. (a) Baseline (b) Larger BEV size (c) Applying RC-Sampling (d) Applying In-Box Label
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent DiffusionWentao Qu, Guofeng Mei, Jing Wang, Yujiao Wu et al.AAAI 2026 · 6 citations
- RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy PredictionBaojie Fan, Xiaotian Li, Yuhan Zhou, Yuyu Jiang et al.ICCV 2025 · 1 citation
- DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous DrivingHongbin Lin, Yiming Yang, Chaoda Zheng, Yifan Zhang et al.AAAI 2026 · 1 citation
- Towards 3D Object-Centric Feature Learning for Semantic Scene CompletionWeihua Wang, Yubo Cui, Xiangru Lin, Zhiheng Li et al.AAAI 2026
- STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object DetectionHuijie Fan, Pengrui Huang, Qiang Wang, Baojie Fan et al.CVPR 2026
Builds on14
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li et al.ICCV 2023 · 399 citations
- Not All Points Are Equal: Learning Highly Efficient Point-based Detectors for 3D LiDAR Point CloudsYifan Zhang, Qingyong Hu, Guoquan Xu, Yanxin Ma et al.CVPR 2022 · 376 citations
Related papers
- EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionYongjin Lee, Hyeon Mun Jeong, Yurim Jeon, Sanghyun KimICCV 2025 · 5 citations
- A Versatile Multi-View Framework for LiDAR-based 3D Object Detection with Guidance from Panoptic SegmentationHamidreza Fazlali, Yixuan Xu, Yuan Ren, Bingbing LiuCVPR 2022 · 23 citations
- DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge DistillationZeyu Wang, Dingwen Li, Chenxu Luo, Cihang Xie et al.ICCV 2023 · 65 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang et al.ICLR 2023 · 28 citations
