UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models Prior
Yao Wu, Mingwei Xing, Yachao Zhang, Xiaotong Luo, Yuan Xie, Yanyun Qu
摘要
3D semantic segmentation using an adapting model trained from a source domain with or without accessing unlabeled target-domain data is the fundamental task in computer vision, containing domain adaptation and domain generalization. The essence of simultaneously solving cross-domain tasks is to enhance the generalizability of the encoder. In light of this, we propose a groundbreaking universal method with the help of off-the-shelf Visual Foundation Models (VFMs) to boost the adaptability and generalizability of cross-domain 3D semantic segmentation, dubbed UniDSeg . Our method explores the VFMs prior and how to harness them, aiming to inherit the recognition ability of VFMs. Specifically, this method introduces layer-wise learnable blocks to the VFMs, which hinges on alternately learning two representations during training: (i) Learning visual prompt. The 3D-to-2D transitional prior and task-shared knowledge is captured from the prompt space, and then (ii) Learning deep query. Spatial Tunability is constructed to the representation of distinct instances driven by prompts in the query space. Integrating these representations into a cross-modal learning framework, UniDSeg efficiently mitigates the domain gap between 2D and 3D modalities, achieving unified cross-domain 3D semantic segmentation. Extensive experiments demonstrate the effectiveness of our method across widely recognized tasks and datasets, all achieving superior performance over state-of-the-art methods. Remarkably, UniDSeg achieves 57.5%/54.4% mIoU on “A2D2/sKITTI” for domain adaptive/generalized tasks. Code is available at https://github.com/Barcaaaa/UniDSeg .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion ModelsFan Li, Xuan Wang, Xuanbin Wang, Zhaoxiang Zhang 等NeurIPS 2025 · 被引用 4 次
- Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic SegmentationXuweiyi Chen, Wentao Zhou, Aruni RoyChowdhury, Zezhou ChengICLR 2026 · 被引用 4 次
- Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability PerspectiveJiahao Li, Yang Lu, Yachao Zhang, Yong Xie 等AAAI 2026 · 被引用 3 次
- UniDxMD: Towards Unified Representation for Cross-Modal Unsupervised Domain Adaptation in 3D Semantic SegmentationZhengyin Liang, Hui Yin, Min Liang, Qianqian Du 等ICCV 2025 · 被引用 2 次
- PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous DrivingYining Pan, Shijie Li, Yuchen Wu, Xulei Yang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- Transfer Learning from Synthetic to Real LiDAR Point Cloud for Semantic SegmentationAoran Xiao, Jiaxing Huang, Dayan Guan, Fangneng Zhan 等AAAI 2022 · 被引用 144 次
- Generalize then Adapt: Source-Free Domain Adaptive Semantic SegmentationJogendra Nath Kundu, Akshay R. Kulkarni, Amit Singh, Varun Jampani 等ICCV 2021 · 被引用 143 次
相关 Paper
- Unlocking 3D Affordance Segmentation with 2D Semantic KnowledgeYu Huang, Zelin Peng, Changsong Wen, Xiaokang Yang 等CVPR 2026 · 被引用 3 次
- Unleashing the Power of Visual Foundation Models for Generalizable Semantic SegmentationPeiyuan Tang, Xiaodong Zhang, Chunze Yang, Haoran Yuan 等AAAI 2025 · 被引用 3 次
- AdaCo: Overcoming Visual Foundation Model Noise in 3D Semantic Segmentation via Adaptive Label CorrectionPufan Zou, Shijia Zhao, Weijie Huang, Qiming Xia 等AAAI 2025
- Generalizable Knowledge Distillation from Vision Foundation Models for Semantic SegmentationChonghua Lv, Dong Zhao, Shuang Wang, Dou Quan 等CVPR 2026 · 被引用 1 次
- UniVS: Unified and Universal Video Segmentation with Prompts as QueriesMinghan Li, Shuai Li, Xindong Zhang, Lei ZhangCVPR 2024
