XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation
Ziyi Wang, Yanbo Wang, Xumin Yu, Jie Zhou, Jiwen Lu
Abstract
Existing methodologies in open vocabulary 3D semantic segmentation primarily concentrate on establishing a unified feature space encompassing 3D, 2D, and textual modalities. Nevertheless, traditional techniques such as global feature alignment or vision-language model distillation tend to impose only approximate correspondence, struggling notably with delineating fine-grained segmentation boundaries. To address this gap, we propose a more meticulous mask-level alignment between 3D features and the 2D-text embedding space through a cross-modal mask reasoning framework, XMask3D. In our approach, we developed a mask generator based on the denoising UNet from a pre-trained diffusion model, leveraging its capability for precise textual control over dense pixel representations and enhancing the open-world adaptability of the generated masks. We further integrate 3D global features as implicit conditions into the pre-trained 2D denoising UNet, enabling the generation of segmentation masks with additional 3D geometry awareness. Subsequently, the generated 2D masks are employed to align mask-level 3D representations with the vision-language feature space, thereby augmenting the open vocabulary capability of 3D geometry embeddings. Finally, we fuse complementary 2D and 3D mask features, resulting in competitive performance across multiple benchmarks for 3D open vocabulary semantic segmentation. Code is available at https://github.com/wangzy22/XMask3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 506abb2f-604a-4fca-9b0f-e01796a530e7Cited by top-tier papers3
- GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic SegmentationXujing Tao, Chuxin Wang, Yubo Ai, Zhixin Cheng et al.CVPR 2026 · 3 citations
- PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global CurriculumShiqi Zhang, Sha Zhang, Jiajun Deng, Yedong Shen et al.ACM MM 2025 · 2 citations
- GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D SegmentationWeijia Dou, Xu Zhang, Yi Bin, Jian Liu et al.ICLR 2026 · 1 citation
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Identity-Aware Language Gaussian Splatting for Open-Vocabulary 3D Semantic SegmentationSungMin Jang, Wonjun KimICCV 2025 · 1 citation
- Open-Vocabulary 3D Semantic Segmentation with Foundation ModelsLi Jiang, Shaoshuai Shi, Bernt SchieleCVPR 2024
- SAS: Segment Any 3D Scene with Integrated 2D PriorsZhuoyuan Li, Jiahao Lu, Jiacheng Deng, Hanzhi Chang et al.ICCV 2025 · 1 citation
- Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene UnderstandingJinlong Li, Cristiano Saltori, Fabio Poiesi, Nicu SebeCVPR 2025
- DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic SegmentationZiyu Zhao, Xiaoguang Li, Lingjia Shi, Nasrin Imanpour et al.CVPR 2025
