CLIP2UDA: Making Frozen CLIP Reward Unsupervised Domain Adaptation in 3D Semantic Segmentation
Yao Wu, Mingwei Xing, Yachao Zhang, Yuan Xie, Yanyun Qu
Abstract
Multi-modal Unsupervised Domain Adaptation (MM-UDA) for large-scale 3D semantic segmentation involves adapting 2D and 3D models to a target domain without labels, which significantly reduces the labor-intensive annotations. Existing MM-UDA methods have often attempted to mitigate the domain discrepancy by aligning features between the source and target data. However, this implementation falls short when applied to image perception due to the susceptibility of images to environmental changes compared to point clouds. To mitigate this limitation, in this work, we explore the potentials of an off-the-shelf Contrastive Language-Image Pre-training (CLIP) model with rich whilst heterogeneous knowledge. To make CLIP task-specific, we propose a top-performing method, dubbed CLIP2UDA, which makes frozen CLIP reward unsupervised domain adaptation in 3D semantic segmentation. Specifically, CLIP2UDA alternates between two steps during adaptation: (a) Learning task-specific prompt. 2D features response from the visual encoder are employed to initiate the learning of adaptive text prompt of each domain, and (b) Learning multi-modal domain-invariant representations. These representations interact hierarchically in the shared decoder to obtain unified 2D visual predictions. This enhancement allows for effective alignment between the modality-specific 3D and unified feature space via cross-modal mutual learning. Extensive experimental results demonstrate that our method outperforms state-of-the-art competitors in several widely-recognized adaptation scenarios. Code is available at: https://github.com/Barcaaaa/CLIP2UDA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 98ba8f7d-f921-4d2f-bed8-e2b6492dbb04Cited by top-tier papers9
- UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models PriorYao Wu, Mingwei Xing, Yachao Zhang, Xiaotong Luo et al.NeurIPS 2024 · 15 citations
- Differential Contrastive Training for Gaze EstimationLin Zhang, Yi Tian, Xiyun Wang, Wanru Xu et al.ACM MM 2025 · 5 citations
- Decoupled Global-Local Alignment for Improving Compositional UnderstandingXiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu et al.ACM MM 2025 · 4 citations
- UniDxMD: Towards Unified Representation for Cross-Modal Unsupervised Domain Adaptation in 3D Semantic SegmentationZhengyin Liang, Hui Yin, Min Liang, Qianqian Du et al.ICCV 2025 · 2 citations
- BeyondSparse: Facilitating Mamba to Enhance Cross-Domain 3D Semantic Segmentation in Adverse WeatherYao Wu, Mingwei Xing, Yachao Zhang, Fangyong Wang et al.AAAI 2026 · 1 citation
Related papers
- CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIPRunnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu et al.CVPR 2023
- CLIP2Pose: Frozen CLIP as Semantic Guide for Domain Adaptive Pose EstimationJiawen Li, Fei Jiang, Dandan Zhu, Jinxin Shi et al.AAAI 2026
- Transferring CLIP's Knowledge into Zero-Shot Point Cloud Semantic SegmentationYuanbin Wang, Shaofei Huang, Yulu Gao, Zhen Wang et al.ACM MM 2023 · 17 citations
- CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-TrainingTianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang et al.ICCV 2023 · 220 citations
- All in One: Visual-Description-Guided Unified Point Cloud SegmentationZongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang et al.ICCV 2025 · 1 citation
