Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-View
Shuo Wang, Xinhai Zhao, Hai-Ming Xu, Zehui Chen, Dameng Yu, Jiahao Chang, Zhen Yang, Feng Zhao
摘要
Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D object detection have been continuously proposed, most of them may risk drastic performance degradation when the domain of input images differs from that of training. In this paper, we first analyze the causes of the domain gap for the MV3D-Det task. Based on the covariate shift assumption, we find that the gap mainly attributes to the feature distribution of BEV, which is determined by the quality of both depth estimation and 2D image's feature representation. To acquire a robust depth prediction, we propose to decouple the depth estimation from the intrinsic parameters of the camera (i.e. the focal length) through converting the prediction of metric depth to that of scale-invariant depth and perform dynamic perspective augmentation to increase the diversity of the extrinsic parameters (i.e. the camera poses) by utilizing homography. Moreover, we modify the focal length values to create multiple pseudo-domains and construct an adversarial training loss to encourage the feature representation to be more domain-agnostic. Without bells and whistles, our approach, namely DG-BEV, successfully alleviates the performance drop on the unseen target domain without impairing the accuracy of the source domain. Extensive experiments on Waymo, nuScenes, and Lyft, demonstrate the generalization and effectiveness of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous DrivingHao Lu, Tianshuo Xu, Wenzhao Zheng, Yunpeng Zhang 等NeurIPS 2025 · 被引用 26 次
- DVGT: Driving Visual Geometry TransformerSicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu 等CVPR 2026 · 被引用 23 次
- Unified Domain Generalization and Adaptation for Multi-View 3D Object DetectionGyusam Chang, Jiwon Lee, Donghyun Kim, Jinkyu Kim 等NeurIPS 2024 · 被引用 19 次
- Towards Universal LiDAR-Based 3D Object Detection by Multi-Domain Knowledge TransferGuile Wu, Tongtong Cao, Bingbing Liu, Xingxin Chen 等ICCV 2023 · 被引用 7 次
- Towards Generalizable Multi-Camera 3D Object Detection via Perspective RenderingHao Lu, Yunpeng Zhang, Guoqing Wang, Qing Lian 等AAAI 2025 · 被引用 5 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li 等ICCV 2021 · 被引用 404 次
相关 Paper
- Geometry-Guided Domain Generalization for Monocular 3D Object DetectionFan Yang, Hui Chen, Yuwei He, Sicheng Zhao 等AAAI 2024 · 被引用 12 次
- CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object DetectionGyusam Chang, Wonseok Roh, Sujin Jang, Dongwook Lee 等AAAI 2024 · 被引用 8 次
- Learning Transferable Features for Point Cloud Detection via 3D Contrastive Co-trainingYihan Zeng, Chunwei Wang, Yunbo Wang, Hang Xu 等NeurIPS 2021 · 被引用 36 次
- CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object DetectionZhaonian Kuang, Rui Ding, Haotian Wang, Xinhu Zheng 等CVPR 2026 · 被引用 2 次
- DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge DistillationZeyu Wang, Dingwen Li, Chenxu Luo, Cihang Xie 等ICCV 2023 · 被引用 65 次
