Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
Yan Wang, Baoxiong Jia, Ziyu Zhu, Siyuan Huang
Abstract
Open-vocabulary 3D scene understanding is pivotal for enhancing physical intelligence, as it enables embodied agents to interpret and interact dynamically within real-world environments. This paper introduces MPEC, a novel Masked Point-Entity Contrastive learning method for open-vocabulary 3D semantic segmentation that leverages both 3D entity-language alignment and point-entity consistency across different point cloud views to foster entity-specific feature representations. Our method improves semantic discrimination and enhances the differentiation of unique instances, achieving state-of-the-art results on ScanNet for open-vocabulary 3D semantic segmentation and demonstrating superior zero-shot scene understanding capabilities. Extensive fine-tuning experiments on 8 datasets, spanning from low-level perception to high-level reasoning tasks, showcase the potential of learned 3D features, driving consistent performance gains across varied 3D scene understanding tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a350d25-55a9-4405-85bf-0b3f98acb64cCited by top-tier papers7
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D DataZhe Zhu, Le Wan, Rui Xu, Yiheng Zhang et al.ICLR 2026 · 15 citations
- GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic SegmentationXujing Tao, Chuxin Wang, Yubo Ai, Zhixin Cheng et al.CVPR 2026 · 3 citations
- PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud LearningXinxing Yu, Ajian Liu, Sunyuan Qiang, Hui Ma et al.CVPR 2026 · 1 citation
- GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D SegmentationWeijia Dou, Xu Zhang, Yi Bin, Jian Liu et al.ICLR 2026 · 1 citation
- FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object SegmentationZihui Zhang, Zhixuan Sun, Yafei YANG, Jinxi Li et al.ICML 2026
Builds on56
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
Related papers
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys et al.NeurIPS 2023 · 389 citations
- Open-Vocabulary 3D Semantic Segmentation with Foundation ModelsLi Jiang, Shaoshuai Shi, Bernt SchieleCVPR 2024
- RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene UnderstandingJihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang et al.CVPR 2024 · 49 citations
- OV3D-CG: Open-Vocabulary 3D Instance Segmentation with Contextual GuidanceMingquan Zhou, Chen He, Ruiping Wang, Xilin ChenICCV 2025 · 1 citation
- Identity-Aware Language Gaussian Splatting for Open-Vocabulary 3D Semantic SegmentationSungMin Jang, Wonjun KimICCV 2025 · 1 citation
