PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
Yining Pan, Shijie Li, Yuchen Wu, Xulei Yang, Na Zhao
Abstract
This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve generalization under domain shifts commonly encountered in real-world autonomous driving. A straightforward solution is to employ a pseudo-labeling strategy, which is widely used in UDA to generate supervision for unlabeled target data, combined with an mm-3DPS backbone. However, existing supervised mm-3DPS methods rely heavily on strong cross-modal complementarity between LiDAR and RGB inputs, making them fragile under domain shifts where one modality degrades (e.g., poor lighting or adverse weather). Moreover, conventional pseudo-labeling typically retains only high-confidence regions, leading to fragmented masks and incomplete object supervision, which are issues particularly detrimental to panoptic segmentation. To address these challenges, we propose PanDA, the first UDA framework specifically designed for multimodal 3D panoptic segmentation. To improve robustness against single-sensor degradation, we introduce an asymmetric multimodal augmentation that selectively drops regions to simulate domain shifts and improve robust representation learning. To enhance pseudo-label completeness and reliability, we further develop a dual-expert pseudo-label refinement module that extracts domain-invariant priors from both 2D and 3D modalities. Extensive experiments across diverse domain shifts, spanning time, weather, location, and sensor variations, significantly surpass state-of-the-art UDA baselines for 3D semantic segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05841619-7ee0-4ba9-a116-25d4ce059444Cited by top-tier papers1
Ask how each one uses itBuilds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- LidarMultiNet: Towards a Unified Multi-Task Network for LiDAR PerceptionDongqiangzi Ye, Zixiang Zhou, Weijia Chen, Yufei Xie et al.AAAI 2023 · 108 citations
Related papers
- LiDAR-UDA: Self-ensembling Through Time for Unsupervised LiDAR Domain AdaptationAmirreza Shaban, Joonho Lee, Sanghun Jung, Xiangyun Meng et al.ICCV 2023 · 20 citations
- Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-Dataset 3D Object DetectionZhanwei Zhang, Minghao Chen, Shuai Xiao, Liang Peng et al.CVPR 2024 · 10 citations
- SPG: Unsupervised Domain Adaptation for 3D Object Detection via Semantic Point GenerationQiangeng Xu, Yin Zhou, Weiyue Wang, Charles R. Qi et al.ICCV 2021 · 172 citations
- MSeg3D: Multi-Modal 3D Semantic Segmentation for Autonomous DrivingJiale Li, Hang Dai, Hao Han, Yong DingCVPR 2023
- Denoise and Align: Towards Source-Free UDA for Robust Panoramic Semantic SegmentationYaowen Chang, Zhen Cao, Xu Zheng, Xiaoxin Mi et al.CVPR 2026 · 4 citations
