Query Refinement Transformer for 3D Instance Segmentation
Jiahao Lu, Jiacheng Deng, Chuxin Wang, Jianfeng He, Tianzhu Zhang
Abstract
3D instance segmentation aims to predict a set of object instances in a scene and represent them as binary foreground masks with corresponding semantic labels. However, object instances are diverse in shape and category, and point clouds are usually sparse, unordered, and irregular, which leads to a query sampling dilemma. Besides, noise background queries interfere with proper scene perception and accurate instance segmentation. To address the above issues, we propose the Query Refinement Transformer termed QueryFormer. The key to our approach is to exploit a query initialization module to optimize the initialization process for the query distribution with a high coverage and low repetition rate. Additionally, we design an affiliated transformer decoder that suppresses the interference of noise background queries and helps the foreground queries focus on instance discriminative parts to predict final segmentation results. Extensive experiments on Scan-NetV2 and S3DIS datasets show that our QueryFormer can surpass state-of-the-art 3D instance segmentation methods. * Corresponding Author GT Mask3D Mask3D Ours Ours C B GT Mask3D Mask3D Ours Ours w/o denoising module w denoising module GT A w/o denoising module w denoising module
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers31
- DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene RenderingJiahao Lu, Jiacheng Deng, Ruijie Zhu, Yanzhe Liang et al.NeurIPS 2024 · 37 citations
- A Unified Framework for 3D Scene UnderstandingWei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou et al.NeurIPS 2024 · 25 citations
- SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature AlignmentQi Xu, Dongxu Wei, Lingzhe Zhao, Wenpu Li et al.NeurIPS 2025 · 19 citations
- ForestFormer3D: A Unified Framework for End-to-End Segmentation of Forest LiDAR 3D Point CloudsBinbin Xiang, Maciej Wielgosz, Stefano Puliti, Kamil Král et al.ICCV 2025 · 12 citations
- Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical RepresentationSangyun Shin, Kaichen Zhou, Madhu Vankadari, Andrew Markham et al.CVPR 2024 · 12 citations
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
Related papers
- 3D Instance Segmentation via Enhanced Spatial and Semantic SupervisionSalwa K. Al Khatib, Mohamed El Amine Boudjoghra, Jean Lahoud, Fahad Shahbaz KhanICCV 2023 · 10 citations
- Superpoint Transformer for 3D Scene Instance SegmentationJiahao Sun, Chunmei Qing, Junpeng Tan, Xiangmin XuAAAI 2023 · 181 citations
- CompetitorFormer: Mitigating Query Conflicts for 3D Instance Segmentation via Competitive StrategyDuanchu Wang, Junjie Yang, Haoran Gong, Jing Liu et al.CVPR 2026
- Relation3D : Enhancing Relation Modeling for Point Cloud Instance SegmentationJiahao Lu, Jiacheng DengCVPR 2025
- MSTA3D: Multi-scale Twin-attention for 3D Instance SegmentationDuc Dang Trung Tran, Byeongkeun Kang, Yeejin LeeACM MM 2024 · 6 citations
