Bridging the Domain Gap: Self-Supervised 3D Scene Understanding with Foundation Models
Zhimin Chen, Longlong Jing, Yingwei Li, Bing Li
Abstract
Foundation models have achieved remarkable results in 2D and language tasks like image segmentation, object detection, and visual-language understanding. However, their potential to enrich 3D scene representation learning is largely untapped due to the existence of the domain gap. In this work, we propose an innovative methodology called Bridge3D to address this gap by pre-training 3D models using features, semantic masks, and captions sourced from foundation models. Specifically, our method employs semantic masks from foundation models to guide the masking and reconstruction process for the masked autoencoder, enabling more focused attention on foreground representations. Moreover, we bridge the 3D-text gap at the scene level using image captioning foundation models, thereby facilitating scene-level knowledge distillation. We further extend this bridging effort by introducing an innovative object-level knowledge distillation method that harnesses highly accurate object-level masks and semantic text data from foundation models. Our methodology significantly surpasses the performance of existing state-of-the-art methods in 3D object detection and semantic segmentation tasks. For instance, on the ScanNet dataset, Bridge3D improves the baseline by a notable margin of 6.3%. Code will be available at: https://github.com/Zhimin-C/Bridge3D
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e66195ff-dcd3-42f7-8b0c-85453dbd5c72Cited by top-tier papers10
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial RepresentationsYujia Zhang, Xiaoyang Wu, Yixing Lao, Chengyao Wang et al.NeurIPS 2025 · 47 citations
- CVRecon: Rethinking 3D Geometric Feature Learning For Neural ReconstructionZiyue Feng, Liang Yang, Pengsheng Guo, Bing LiICCV 2023 · 28 citations
- SAM-Guided Masked Token Prediction for 3D Scene UnderstandingZhimin Chen, Liang Yang, Yingwei Li, Longlong Jing et al.NeurIPS 2024 · 12 citations
- CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ SegmentationXinlei Yu, Changmiao Wang, Hui Jin, Ahmed Elazab et al.ACM MM 2025 · 3 citations
- Semantic Foam: Unifying Spatial and Semantic Scene DecompositionAmr Sharafeldin, Aryan Mikaeili, Thomas Walker, Shrisudhan Govindarajan et al.CVPR 2026 · 1 citation
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D SegmentationJunha Lee, Chunghyun Park, Jaesung Choe, Yu-Chiang Frank Wang et al.CVPR 2025
- Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?Runpei Dong, Zekun Qi, Linfeng Zhang, Junbo Zhang et al.ICLR 2023 · 21 citations
- From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMsAng Cao, Sergio Arnaud, Oleksandr Maksymets, Jianing Yang et al.ICML 2025
- 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene UnderstandingXiaohu Huang, Jingjing Wu, Qunyi Xie, Kai HanNeurIPS 2025 · 11 citations
- XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic SegmentationZiyi Wang, Yanbo Wang, Xumin Yu, Jie Zhou et al.NeurIPS 2024 · 7 citations
