OneFormer3D: One Transformer for Unified Point Cloud Segmentation
Maxim Kolodiazhnyi, Anna Vorontsova, Anton Konushin, Danila Rukhovich
Abstract
Semantic, instance, and panoptic segmentation of 3D point clouds have been addressed using task-specific models of distinct design. Thereby, the similarity of all segmentation tasks and the implicit relationship between them have not been utilized effectively. This paper presents a unified, simple, and effective model addressing all these tasks jointly. The model, named OneFormer3D, performs instance and semantic segmentation consistently, using a group of learnable kernels, where each kernel is responsible for generating a mask for either an instance or a semantic category. These kernels are trained with a transformerbased decoder with unified instance and semantic queries passed as an input. Such a design enables training a model end-to-end in a single run, so that it achieves top performance on all three segmentation tasks simultaneously. Specifically, our OneFormer3D ranks 1 st and sets a new state-of-the-art (+2.1 mAP 50 ) in the ScanNet test leaderboard. We also demonstrate the state-of-the-art results in semantic, instance, and panoptic segmentation of ScanNet (+21 PQ), ScanNet200 (+3.8 mAP 50 ), and S3DIS (+0.8 mIoU) datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d97d396d-3f7d-4357-9777-a8243752ea82Cited by top-tier papers59
- Chat-Scene: Bridging 3D Scene and Large Language Models with Object IdentifiersHaifeng Huang, Yilun Chen, Zehan Wang, Rongjie Huang et al.NeurIPS 2024 · 230 citations
- A Unified Framework for 3D Scene UnderstandingWei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou et al.NeurIPS 2024 · 25 citations
- SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature AlignmentQi Xu, Dongxu Wei, Lingzhe Zhao, Wenpu Li et al.NeurIPS 2025 · 19 citations
- ForestFormer3D: A Unified Framework for End-to-End Segmentation of Forest LiDAR 3D Point CloudsBinbin Xiang, Maciej Wielgosz, Stefano Puliti, Kamil Král et al.ICCV 2025 · 12 citations
- RefMask3D: Language-Guided Transformer for 3D Referring SegmentationShuting He, Henghui DingACM MM 2024 · 12 citations
Builds on27
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai et al.NeurIPS 2022 · 1,270 citations
Related papers
- Unified 3D Segmenter As Prototypical ClassifiersZheyun Qin, Cheng Han, Qifan Wang, Xiushan Nie et al.NeurIPS 2023 · 27 citations
- OneFormer: One Transformer to Rule Universal Image SegmentationJitesh Jain, Jiachen Li, MangTik Chiu, Ali Hassani et al.CVPR 2023
- Superpoint Transformer for 3D Scene Instance SegmentationJiahao Sun, Chunmei Qing, Junpeng Tan, Xiangmin XuAAAI 2023 · 181 citations
- 3D Instance Segmentation via Enhanced Spatial and Semantic SupervisionSalwa K. Al Khatib, Mohamed El Amine Boudjoghra, Jean Lahoud, Fahad Shahbaz KhanICCV 2023 · 10 citations
- PUPS: Point Cloud Unified Panoptic SegmentationShihao Su, Jianyun Xu, Huanyu Wang, Zhenwei Miao et al.AAAI 2023 · 30 citations
