BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic Segmentation
Miaoyu Li, Yachao Zhang, Xu Ma, Yanyun Qu, Yun Fu
Abstract
Cross-modal Unsupervised Domain Adaptation aims to exploit the complementarity of 2D-3D data to overcome the lack of annotation in an unknown domain. However, the training of these methods relies on access to target samples, meaning the trained model only works in a specific target domain. In light of this, we propose cross-modal learning under bird’s-eye view for Domain Generalization (DG) of 3D semantic segmentation, called BEV-DG. DG is more challenging because the model cannot access the target domain during training, meaning it needs to rely on cross-modal learning to alleviate the domain gap. Since 3D semantic segmentation requires the classification of each point, existing cross-modal learning is directly conducted point-to-point, which is sensitive to the misalignment in projections between pixels and points. To this end, our approach aims to optimize domain-irrelevant representation modeling with the aid of cross-modal learning under bird’s-eye view. We propose BEV-based Area-to-area Fusion (BAF) to conduct cross-modal learning under bird’s-eye view, which has a higher fault tolerance for point-level misalignment. Furthermore, to model domain-irrelevant representations, we propose BEV-driven Domain Contrastive Learning (BDCL) with the help of cross-modal learning under bird’s-eye view. We design three domain generalization settings based on three 3D datasets, and BEV-DG significantly outperforms state-of-the-art competitors with tremendous margins in all settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c697fa32-c7c3-4de7-b048-5e0e63151b5fCited by top-tier papers5
- UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models PriorYao Wu, Mingwei Xing, Yachao Zhang, Xiaotong Luo et al.NeurIPS 2024 · 15 citations
- No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion ModelsFan Li, Xuan Wang, Xuanbin Wang, Zhaoxiang Zhang et al.NeurIPS 2025 · 4 citations
- Adaptive Augmentation-Aware Latent Learning for Robust LiDAR Semantic SegmentationWangkai Li, Zhaoyang Li, Yuwen Pan, Rui Sun et al.ICLR 2026 · 1 citation
- Semi-Supervised Semantic Segmentation via Derivative Label PropagationYuanbin Fu, Xiaojie GuoAAAI 2026 · 1 citation
- ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy PredictionYuheng Zhang, Mengfei Duan, Kunyu Peng, Yuhang Wang et al.CVPR 2026
Builds on18
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Deep Domain-Adversarial Image Generation for Domain GeneralisationKaiyang Zhou, Yongxin Yang, Timothy M. Hospedales, Tao XiangAAAI 2020 · 488 citations
- Episodic Training for Domain GeneralizationDa Li, Jianshu Zhang, Yongxin Yang, Cong Liu et al.ICCV 2019 · 488 citations
- ShellNet: Efficient Point Cloud Convolutional Neural Networks Using Concentric Shells StatisticsZhiyuan Zhang, Binh-Son Hua, Sai-Kit YeungICCV 2019 · 400 citations
- Efficient Domain Generalization via Common-Specific Low-Rank DecompositionVihari Piratla, Praneeth Netrapalli, Sunita SarawagiICML 2020 · 250 citations
Related papers
- CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object DetectionGyusam Chang, Wonseok Roh, Sujin Jang, Dongwook Lee et al.AAAI 2024 · 8 citations
- Cross-modal Unsupervised Domain Adaptation for 3D Semantic Segmentation via Bidirectional Fusion-then-DistillationYao Wu, Mingwei Xing, Yachao Zhang, Yuan Xie et al.ACM MM 2023 · 21 citations
- Cross-Modal Contrastive Learning for Domain Adaptation in 3D Semantic SegmentationBowei Xing, Xianghua Ying, Ruibin Wang, Jinfa Yang et al.AAAI 2023 · 23 citations
- Self-supervised Exclusive Learning for 3D Segmentation with Cross-Modal Unsupervised Domain AdaptationYachao Zhang, Miaoyu Li, Yuan Xie, Cuihua Li et al.ACM MM 2022 · 22 citations
- Learning Transferable Features for Point Cloud Detection via 3D Contrastive Co-trainingYihan Zeng, Chunwei Wang, Yunbo Wang, Hang Xu et al.NeurIPS 2021 · 36 citations
