3D Instance Segmentation via Enhanced Spatial and Semantic Supervision
Salwa K. Al Khatib, Mohamed El Amine Boudjoghra, Jean Lahoud, Fahad Shahbaz Khan
Abstract
3D instance segmentation has recently garnered increased attention. Typical deep learning methods adopt point grouping schemes followed by hand-designed geometric clustering. Inspired by the success of transformers for various 3D tasks, newer hybrid approaches have utilized transformer decoders coupled with convolutional backbones that operate on voxelized scenes. However, due to the nature of sparse feature backbones, the extracted features provided to the transformer decoder are lacking in spatial understanding. Thus, such approaches often predict spatially separate objects as single instances. To this end, we introduce a novel approach for 3D point clouds instance segmentation that addresses the challenge of generating distinct instance masks for objects that share similar appearances but are spatially separated. Our method leverages spatial and semantic supervision with query refinement to improve the performance of hybrid 3D instance segmentation models. Specifically, we provide the transformer block with spatial features to facilitate differentiation between similar object queries and incorporate semantic supervision to enhance prediction accuracy based on object class. Our proposed approach outperforms existing methods on the validation sets of ScanNet V2 and ScanNet200 datasets, establishing a new state-of-the-art for this task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 881cb6bf-e7d4-48dc-a7b6-2739ecd4b0ecCited by top-tier papers4
- Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical RepresentationSangyun Shin, Kaichen Zhou, Madhu Vankadari, Andrew Markham et al.CVPR 2024 · 12 citations
- Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance SegmentationMohamed El Amine Boudjoghra, Angela Dai, Jean Lahoud, Hisham Cholakkal et al.ICLR 2025 · 3 citations
- SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D FeaturesJinyuan Qu, Hongyang Li, Xingyu Chen, Shilong Liu et al.AAAI 2026 · 2 citations
- All in One: Visual-Description-Guided Unified Point Cloud SegmentationZongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang et al.ICCV 2025 · 1 citation
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- Stratified Transformer for 3D Point Cloud SegmentationXin Lai, Jianhui Liu, Li Jiang, Liwei Wang et al.CVPR 2022 · 494 citations
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu et al.ICCV 2021 · 368 citations
Related papers
- Query Refinement Transformer for 3D Instance SegmentationJiahao Lu, Jiacheng Deng, Chuxin Wang, Jianfeng He et al.ICCV 2023 · 56 citations
- Superpoint Transformer for 3D Scene Instance SegmentationJiahao Sun, Chunmei Qing, Junpeng Tan, Xiangmin XuAAAI 2023 · 181 citations
- MSTA3D: Multi-scale Twin-attention for 3D Instance SegmentationDuc Dang Trung Tran, Byeongkeun Kang, Yeejin LeeACM MM 2024 · 6 citations
- Relation3D : Enhancing Relation Modeling for Point Cloud Instance SegmentationJiahao Lu, Jiacheng DengCVPR 2025
- OneFormer3D: One Transformer for Unified Point Cloud SegmentationMaxim Kolodiazhnyi, Anna Vorontsova, Anton Konushin, Danila RukhovichCVPR 2024
