Superpoint Transformer for 3D Scene Instance Segmentation
Jiahao Sun, Chunmei Qing, Junpeng Tan, Xiangmin Xu
Abstract
Most existing methods realize 3D instance segmentation by extending those models used for 3D object detection or 3D semantic segmentation. However, these non-straightforward methods suffer from two drawbacks: 1) Imprecise bounding boxes or unsatisfactory semantic predictions limit the performance of the overall 3D instance segmentation framework. 2) Existing methods require a time-consuming intermediate step of aggregation. To address these issues, this paper proposes a novel end-to-end 3D instance segmentation method based on Superpoint Transformer, named as SPFormer. It groups potential features from point clouds into superpoints, and directly predicts instances through query vectors without relying on the results of object detection or semantic segmentation. The key step in this framework is a novel query decoder with transformers that can capture the instance information through the superpoint cross-attention mechanism and generate the superpoint masks of the instances. Through bipartite matching based on superpoint masks, SPFormer can implement the network training without the intermediate aggregation step, which accelerates the network. Extensive experiments on ScanNetv2 and S3DIS benchmarks verify that our method is concise yet efficient. Notably, SPFormer exceeds compared state-of-the-art methods by 4.3% on Scan-Netv2 hidden test set in terms of mAP and keeps fast inference speed (247ms per frame) simultaneously. Code is available at https://github.com/sunjiahao1999/SPFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 441b28a1-1a3a-4dad-af27-b033e2d0975fCited by top-tier papers58
- Point Cloud Mamba: Point Cloud Learning via State Space ModelTao Zhang, Haobo Yuan, Lu Qi, Jiangning Zhang et al.AAAI 2025 · 110 citations
- Mask-Attention-Free Transformer for 3D Instance SegmentationXin Lai, Yuhui Yuan, Ruihang Chu, Yukang Chen et al.ICCV 2023 · 53 citations
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidancePhuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan et al.CVPR 2024 · 45 citations
- 3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression SegmentationChangli Wu, Yiwei Ma, Qi Chen, Haowei Wang et al.AAAI 2024 · 40 citations
- AGILE3D: Attention Guided Interactive Multi-object 3D SegmentationYuanwen Yue, Sabarinath Mahadevan, Jonas Schult, Francis Engelmann et al.ICLR 2024 · 36 citations
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li et al.ICCV 2021 · 331 citations
- SoftGroup for 3D Instance Segmentation on Point CloudsThang Vu, Kookhoi Kim, Tung Minh Luu, Thanh Xuan Nguyen et al.CVPR 2022 · 251 citations
Related papers
- Query Refinement Transformer for 3D Instance SegmentationJiahao Lu, Jiacheng Deng, Chuxin Wang, Jianfeng He et al.ICCV 2023 · 56 citations
- Relation3D : Enhancing Relation Modeling for Point Cloud Instance SegmentationJiahao Lu, Jiacheng DengCVPR 2025
- MSTA3D: Multi-scale Twin-attention for 3D Instance SegmentationDuc Dang Trung Tran, Byeongkeun Kang, Yeejin LeeACM MM 2024 · 6 citations
- 3D Instance Segmentation via Enhanced Spatial and Semantic SupervisionSalwa K. Al Khatib, Mohamed El Amine Boudjoghra, Jean Lahoud, Fahad Shahbaz KhanICCV 2023 · 10 citations
- PointGroup: Dual-Set Point Grouping for 3D Instance SegmentationLi Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu et al.CVPR 2020
