Geometry-Guided Domain Generalization for Monocular 3D Object Detection
Fan Yang, Hui Chen, Yuwei He, Sicheng Zhao, Chenghao Zhang, Kai Ni, Guiguang Ding
Abstract
Monocular 3D object detection (M3OD) is important for autonomous driving. However, existing deep learning-based methods easily suffer from performance degradation in real-world scenarios due to the substantial domain gap between training and testing. M3OD's domain gaps are complex, including camera intrinsic parameters, extrinsic parameters, image appearance, etc. Existing works primarily focus on the domain gaps of camera intrinsic parameters, ignoring other key factors. Moreover, at the feature level, conventional domain invariant learning methods generally cause the negative transfer issue, due to the ignorance of dependency between geometry tasks and domains. To tackle these issues, in this paper, we propose MonoGDG, a geometry-guided domain generalization framework for M3OD, which effectively addresses the domain gap at both camera and feature levels. Specifically, MonoGDG consists of two major components. One is geometry-based image reprojection, which mitigates the impact of camera discrepancy by unifying intrinsic parameters, randomizing camera orientations, and unifying the field of view range. The other is geometry-dependent feature disentanglement, which overcomes the negative transfer problems by incorporating domain-shared and domain-specific features. Additionally, we leverage a depth-disentangled domain discriminator and a domain-aware geometry regression attention mechanism to account for the geometry-domain dependency. Extensive experiments on multiple autonomous driving benchmarks demonstrate that our method achieves state-of-the-art performance in domain generalization for M3OD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d70ed67-41d6-46ad-926f-dee0b1723d1bCited by top-tier papers4
- CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object DetectionZhaonian Kuang, Rui Ding, Haotian Wang, Xinhu Zheng et al.CVPR 2026 · 2 citations
- Enhancing Generalizability via Utilization of Unlabeled Data for Occupancy PerceptionRuihang Li, Tao Li, Shanding Ye, Kaikai Xiao et al.AAAI 2025 · 1 citation
- Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object DetectionJunshu Zhang, Sicheng Zhao, Xin Zhao, Fan Yang et al.CVPR 2026
- HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility EvaluatorFan Yang, Ru Zhen, Jianing Wang, Yanhao Zhang et al.CVPR 2025
Builds on8
- Geometry Uncertainty Projection Network for Monocular 3D Object DetectionYan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang et al.ICCV 2021 · 294 citations
- Geometry-based Distance Decomposition for Monocular 3D Object DetectionXuepeng Shi, Qi Ye, Xiaozhi Chen, Chuangrong Chen et al.ICCV 2021 · 169 citations
- Physics-Based Rendering for Improving Robustness to RainShirsendu Sukanta Halder, Jean-François Lalonde, Raoul de CharetteICCV 2019 · 129 citations
- PIT: Position-Invariant Transform for Cross-FoV Domain AdaptationQiqi Gu, Qianyu Zhou, Minghao Xu, Zhengyang Feng et al.ICCV 2021 · 44 citations
- Confidence-based Visual Dispersal for Few-shot Unsupervised Domain AdaptationYizhe Xiong, Hui Chen, Zijia Lin, Sicheng Zhao et al.ICCV 2023 · 13 citations
Related papers
- Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-ViewShuo Wang, Xinhai Zhao, Hai-Ming Xu, Zehui Chen et al.CVPR 2023
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 · 175 citations
- Unsupervised Domain Adaptive 3D Detection with Multi-Level ConsistencyZhipeng Luo, Zhongang Cai, Changqing Zhou, Gongjie Zhang et al.ICCV 2021 · 92 citations
- Difficulty-Aware Label-Guided Denoising for Monocular 3D Object DetectionSoyul Lee, Seungmin Baek, Dongbo MinAAAI 2026
- MonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error PriorsFanqi Pu, Yifan Wang, Jiru Deng, Wenming YangCVPR 2025
