MaskClustering: View Consensus Based Mask Graph Clustering for Open-Vocabulary 3D Instance Segmentation
Mi Yan, Jiazhao Zhang, Yan Zhu, He Wang
摘要
Open-vocabulary 3D instance segmentation is cuttingedge for its ability to segment 3D instances without prede-fined categories. However, progress in 3D lags behind its 2D counterpart due to limited annotated 3D data. To ad-dress this, recent works first generate 2D open-vocabulary masks through 2D models and then merge them into 3D instances based on metrics calculated between two neigh-boring frames. In contrast to these local metrics, we pro-pose a novel metric, view consensus rate, to enhance the utilization of multi-view observations. The key insight is that two 2D masks should be deemed part of the same 3D instance if a significant number of other 2D masks from different views contain both these two masks. Using this metric as edge weight, we construct a global mask graph where each mask is a node. Through iterative clustering of masks showing high view consensus, we generate a series of clusters, each representing a distinct 3D instance. Notably, our model is training-free. Through extensive ex-periments on publicly available datasets, including Scan-Net++, ScanNet200 and MatterPo rt3D, we demonstrate that our method achieves state-of-the-art performance in open-vocabulary 3D instance segmentation. Our project page is at https://pku-epic.github.ioIMaskClustering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidancePhuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan 等CVPR 2024 · 被引用 45 次
- SAI3D: Segment any Instance in 3D ScenesYingda Yin, Yuzheng Liu, Yang Xiao, Daniel Cohen-Or 等CVPR 2024 · 被引用 32 次
- MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning SegmentationJiaxin Huang, Runnan Chen, Ziwen Li, Zhengqing Gao 等NeurIPS 2025 · 被引用 18 次
- Intrinsic Image Fusion for Multi-View 3D Material ReconstructionPeter Kocsis, Lukas Höllein, Matthias NießnerCVPR 2026 · 被引用 8 次
- Zoo3D: Zero-Shot 3D Object Detection at Scene LevelAndrey Lemeshko, Bulat Gabdullin, Nikita Drozdov, Anton Konushin 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun 等ICLR 2022 · 被引用 885 次
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 被引用 659 次
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa 等ICCV 2023 · 被引用 620 次
相关 Paper
- OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D ReconstructionHongyang Li, Jinyuan Qu, Lei ZhangICLR 2026 · 被引用 5 次
- MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance SegmentationYibo Zhao, Yigong Zhang, Jin XieCVPR 2026 · 被引用 1 次
- Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance SegmentationMohamed El Amine Boudjoghra, Angela Dai, Jean Lahoud, Hisham Cholakkal 等ICLR 2025 · 被引用 3 次
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys 等NeurIPS 2023 · 被引用 389 次
- OpenM3D: Open Vocabulary Multi-View Indoor 3D Object Detection without Human AnnotationsPeng-Hao Hsu, Ke Zhang, Fu-En Wang, Tao Tu 等ICCV 2025 · 被引用 3 次
