A Unified Framework for 3D Scene Understanding
Wei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou, Dingkang Liang, Xiang Bai
Abstract
We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are typically tailored to a specific task, limiting their understanding of 3D scenes to a task-specific perspective. In contrast, the proposed method unifies six tasks into unified representations processed by the same Transformer. It facilitates inter-task knowledge sharing, thereby promoting comprehensive 3D scene understanding. To take advantage of multi-task unification, we enhance performance by establishing explicit inter-task associations. Specifically, we design knowledge distillation and contrastive learning methods to transfer task-specific knowledge across different tasks. Experiments on three benchmarks, including ScanNet20, ScanRefer, and ScanNet200, demonstrate that the UniSeg3D consistently outperforms current SOTA methods, even those specialized for individual tasks. We hope UniSeg3D can serve as a solid unified baseline and inspire future work. Code and models are available at https://github.com/dk-liang/UniSeg3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5172fb4a-c6b4-4166-b158-0a7e056315eaCited by top-tier papers8
- SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature AlignmentQi Xu, Dongxu Wei, Lingzhe Zhao, Wenpu Li et al.NeurIPS 2025 · 19 citations
- RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression SegmentationChangli Wu, Qi Chen, Jiayi Ji, Haowei Wang et al.NeurIPS 2024 · 16 citations
- OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene UnderstandingYoujun Zhao, Jiaying Lin, Shuquan Ye, Qianshi Pang et al.AAAI 2026 · 5 citations
- IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression SegmentationQi Chen, Changli Wu, Jiayi Ji, Yiwei Ma et al.AAAI 2025 · 5 citations
- PointTPA: Dynamic Network Parameter Adaptation for 3D Scene UnderstandingSiyuan Liu, Chaoqun Zheng, Xin Zhou, Tianrui Feng et al.CVPR 2026 · 3 citations
Builds on46
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai et al.NeurIPS 2022 · 1,270 citations
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu et al.NeurIPS 2022 · 924 citations
- Stratified Transformer for 3D Point Cloud SegmentationXin Lai, Jianhui Liu, Li Jiang, Liwei Wang et al.CVPR 2022 · 494 citations
Related papers
- OneFormer3D: One Transformer for Unified Point Cloud SegmentationMaxim Kolodiazhnyi, Anna Vorontsova, Anton Konushin, Danila RukhovichCVPR 2024
- EPS3D: End-to-End Feed-Forward 3D Panoptic SegmentationRunsong Zhu, Jiaxin GUO, Xiaoyang Guo, Zhengzhe Liu et al.ICML 2026 · 3 citations
- Unified 3D Segmenter As Prototypical ClassifiersZheyun Qin, Cheng Han, Qifan Wang, Xiushan Nie et al.NeurIPS 2023 · 27 citations
- Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View ImagesXiangyu Sun, Haoyi Jiang, Liu Liu, Seungtae Nam et al.CVPR 2026 · 28 citations
- DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation ModelXiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan et al.NeurIPS 2023 · 40 citations
