Towards Generalizable Multi-Camera 3D Object Detection via Perspective Rendering
Hao Lu, Yunpeng Zhang, Guoqing Wang, Qing Lian, Dalong Du, Ying-Cong Chen
Abstract
Detecting objects in 3D space using multiple cameras, known as Multi-Camera 3D Object Detection (MC3D-Det), has gained prominence with the advent of bird'seye view (BEV) approaches. However, these methods often struggle when faced with unfamiliar testing environments due to the lack of diverse training data encompassing various viewpoints and environments. To address this, we propose a novel method that aligns 3D detection with 2D camera plane results, ensuring consistent and accurate detections. Our framework, anchored in perspective debiasing, helps the learning of features resilient to domain shifts. In our approach, we render diverse view maps from BEV features and rectify the perspective bias of these maps, leveraging implicit foreground volumes to bridge the camera and BEV planes. This two-step process promotes the learning of perspective-and context-independent features, crucial for accurate object detection across varying viewpoints, camera parameters and environment conditions. Notably, our modelagnostic approach preserves the original network structure without incurring additional inference costs, facilitating seamless integration across various models and simplifying deployment. Furthermore, we also show our approach achieves satisfactory results in real data when trained only with virtual datasets, eliminating the need for real scene annotations. Experimental results on both Domain Generalization (DG) and Unsupervised Domain Adaptation (UDA) clearly demonstrate its effectiveness. Our code will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18a13db8-38ad-4fff-b51f-392b0600410cCited by top-tier papers4
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal OverheadChaojun Ni, Chen Cheng, Xiaofeng Wang, Zheng Zhu et al.CVPR 2026 · 23 citations
- CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object DetectionZhaonian Kuang, Rui Ding, Haotian Wang, Xinhu Zheng et al.CVPR 2026 · 2 citations
- Generalizable Multi-Camera 3D Object Detection from a Single Source via Fourier Cross-View LearningXue Zhao, Qinying Gu, Xinbing Wang, Chenghu Zhou et al.ICML 2025
- Temporal and Spatial Representation Learning for Multimodal Low-Beam 3D Object DetectionLin Wang, Shiliang Sun, Jing ZhaoAAAI 2026
Builds on14
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain AdaptationTao Sun, Mattia Segù, Janis Postels, Yuxuan Wang et al.CVPR 2022 · 174 citations
- DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous DrivingXiaosong Jia, Yulu Gao, Li Chen, Junchi Yan et al.ICCV 2023 · 154 citations
Related papers
- Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-ViewShuo Wang, Xinhai Zhao, Hai-Ming Xu, Zehui Chen et al.CVPR 2023
- CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object DetectionGyusam Chang, Wonseok Roh, Sujin Jang, Dongwook Lee et al.AAAI 2024 · 8 citations
- UniMODE: Unified Monocular 3D Object DetectionZhuoling Li, Xiaogang Xu, Ser-Nam Lim, Hengshuang ZhaoCVPR 2024
- Unified Domain Generalization and Adaptation for Multi-View 3D Object DetectionGyusam Chang, Jiwon Lee, Donghyun Kim, Jinkyu Kim et al.NeurIPS 2024 · 19 citations
- GPA-3D: Geometry-aware Prototype Alignment for Unsupervised Domain Adaptive 3D Object Detection from Point CloudsZiyu Li, Jingming Guo, Tongtong Cao, Bingbing Liu et al.ICCV 2023 · 19 citations
