Geometry-Guided Domain Generalization for Monocular 3D Object Detection
Fan Yang, Hui Chen, Yuwei He, Sicheng Zhao, Chenghao Zhang, Kai Ni, Guiguang Ding
摘要
Monocular 3D object detection (M3OD) is important for autonomous driving. However, existing deep learning-based methods easily suffer from performance degradation in real-world scenarios due to the substantial domain gap between training and testing. M3OD's domain gaps are complex, including camera intrinsic parameters, extrinsic parameters, image appearance, etc. Existing works primarily focus on the domain gaps of camera intrinsic parameters, ignoring other key factors. Moreover, at the feature level, conventional domain invariant learning methods generally cause the negative transfer issue, due to the ignorance of dependency between geometry tasks and domains. To tackle these issues, in this paper, we propose MonoGDG, a geometry-guided domain generalization framework for M3OD, which effectively addresses the domain gap at both camera and feature levels. Specifically, MonoGDG consists of two major components. One is geometry-based image reprojection, which mitigates the impact of camera discrepancy by unifying intrinsic parameters, randomizing camera orientations, and unifying the field of view range. The other is geometry-dependent feature disentanglement, which overcomes the negative transfer problems by incorporating domain-shared and domain-specific features. Additionally, we leverage a depth-disentangled domain discriminator and a domain-aware geometry regression attention mechanism to account for the geometry-domain dependency. Extensive experiments on multiple autonomous driving benchmarks demonstrate that our method achieves state-of-the-art performance in domain generalization for M3OD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object DetectionZhaonian Kuang, Rui Ding, Haotian Wang, Xinhu Zheng 等CVPR 2026 · 被引用 2 次
- Enhancing Generalizability via Utilization of Unlabeled Data for Occupancy PerceptionRuihang Li, Tao Li, Shanding Ye, Kaikai Xiao 等AAAI 2025 · 被引用 1 次
- Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object DetectionJunshu Zhang, Sicheng Zhao, Xin Zhao, Fan Yang 等CVPR 2026
- HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility EvaluatorFan Yang, Ru Zhen, Jianing Wang, Yanhao Zhang 等CVPR 2025
它引用的顶会 Paper8
- Geometry Uncertainty Projection Network for Monocular 3D Object DetectionYan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang 等ICCV 2021 · 被引用 294 次
- Geometry-based Distance Decomposition for Monocular 3D Object DetectionXuepeng Shi, Qi Ye, Xiaozhi Chen, Chuangrong Chen 等ICCV 2021 · 被引用 169 次
- Physics-Based Rendering for Improving Robustness to RainShirsendu Sukanta Halder, Jean-François Lalonde, Raoul de CharetteICCV 2019 · 被引用 129 次
- PIT: Position-Invariant Transform for Cross-FoV Domain AdaptationQiqi Gu, Qianyu Zhou, Minghao Xu, Zhengyang Feng 等ICCV 2021 · 被引用 44 次
- Confidence-based Visual Dispersal for Few-shot Unsupervised Domain AdaptationYizhe Xiong, Hui Chen, Zijia Lin, Sicheng Zhao 等ICCV 2023 · 被引用 13 次
相关 Paper
- Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-ViewShuo Wang, Xinhai Zhao, Hai-Ming Xu, Zehui Chen 等CVPR 2023
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo 等ICCV 2023 · 被引用 175 次
- Unsupervised Domain Adaptive 3D Detection with Multi-Level ConsistencyZhipeng Luo, Zhongang Cai, Changqing Zhou, Gongjie Zhang 等ICCV 2021 · 被引用 92 次
- Difficulty-Aware Label-Guided Denoising for Monocular 3D Object DetectionSoyul Lee, Seungmin Baek, Dongbo MinAAAI 2026
- MonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error PriorsFanqi Pu, Yifan Wang, Jiru Deng, Wenming YangCVPR 2025
