Open-Vocabulary 3D Semantic Segmentation with Foundation Models
Li Jiang, Shaoshuai Shi, Bernt Schiele
Abstract
In dynamic 3D environments, the ability to recognize a diverse range of objects without the constraints of predefined categories is indispensable for real-world applications. In response to this need, we introduce OV3D, an innovative framework designed for open-vocabulary 3D semantic segmentation. OV3D leverages the broad open-world knowledge embedded in vision and language foundation models to establish a fine-grained correspondence between 3D points and textual entity descriptions. These entity descriptions are enriched with contextual information, enabling a more open and comprehensive understanding. By seamlessly aligning 3D point features with entity text features, OV3D empowers open-vocabulary recognition in the 3D domain, achieving state-of-the-art open-vocabulary semantic segmentation performance across multiple datasets, including ScanNet, Matterport3D, and nuScenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2d87d30-7b40-4d94-ab7e-9a30e4b38434Cited by top-tier papers25
- PartField: Learning 3D Feature Fields for Part Segmentation and BeyondMing-Yu Liu, Mikaela Angelina Uy, Donglai Xiang, Hao Su et al.ICCV 2025 · 103 citations
- MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning SegmentationJiaxin Huang, Runnan Chen, Ziwen Li, Zhengqing Gao et al.NeurIPS 2025 · 18 citations
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D DataZhe Zhu, Le Wan, Rui Xu, Yiheng Zhang et al.ICLR 2026 · 15 citations
- COS3D: Collaborative Open-Vocabulary 3D SegmentationRunsong Zhu, Ka-Hei Hui, Zhengzhe Liu, Qianyi Wu et al.NeurIPS 2025 · 12 citations
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene EncodingYue Li, Qi Ma, Runyi Yang, Mengjiao Ma et al.CVPR 2026 · 10 citations
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- 3D-AVS: LiDAR-based 3D Auto-Vocabulary SegmentationWeijie Wei, Osman Ülger, Fatemeh Karimi Nejadasl, Theo Gevers et al.CVPR 2025
- OV3D-CG: Open-Vocabulary 3D Instance Segmentation with Contextual GuidanceMingquan Zhou, Chen He, Ruiping Wang, Xilin ChenICCV 2025 · 1 citation
- Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D SegmentationJunha Lee, Chunghyun Park, Jaesung Choe, Yu-Chiang Frank Wang et al.CVPR 2025
- OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object DetectionAdrian Chow, Evelien Riddell, Yimu Wang, Sean Sedwards et al.ICCV 2025 · 2 citations
- PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global CurriculumShiqi Zhang, Sha Zhang, Jiajun Deng, Yedong Shen et al.ACM MM 2025 · 2 citations
