GrabS: Generative Embodied Agent for 3D Object Segmentation without Scene Supervision
Zihui Zhang, Yafei Yang, Hongtao Wen, Bo Yang
Abstract
We study the hard problem of 3D object segmentation in complex point clouds without requiring human labels of 3D scenes for supervision. By relying on the similarity of pretrained 2D features or external signals such as motion to group 3D points as objects, existing unsupervised methods are usually limited to identifying simple objects like cars or their segmented objects are often inferior due to the lack of objectness in pretrained features. In this paper, we propose a new twostage pipeline called GrabS. The core concept of our method is to learn generative and discriminative object-centric priors as a foundation from object datasets in the first stage, and then design an embodied agent to learn to discover multiple objects by querying against the pretrained generative priors in the second stage. We extensively evaluate our method on two real-world datasets and a newly created synthetic dataset, demonstrating remarkable segmentation performance, clearly surpassing all existing unsupervised methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene SupervisionJiahao Chen, Zihui Zhang, Yafei Yang, Jinxi Li et al.CVPR 2026 · 1 citation
- FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object SegmentationZihui Zhang, Zhixuan Sun, Yafei YANG, Jinxi Li et al.ICML 2026
- 3D-DLP: Self-supervised 3D Object-centric Scene Representation LearningEllina Zhang, Madhavan Iyengar, Amir Zadeh, Chuan Li et al.ICML 2026
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- LION: Latent Point Diffusion Models for 3D Shape GenerationXiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic et al.NeurIPS 2022 · 752 citations
- Neural Unsigned Distance Fields for Implicit Function LearningJulian Chibane, Aymen Mir, Gerard Pons-MollNeurIPS 2020 · 415 citations
- Superpoint Transformer for 3D Scene Instance SegmentationJiahao Sun, Chunmei Qing, Junpeng Tan, Xiangmin XuAAAI 2023 · 181 citations
Related papers
- OGC: Unsupervised 3D Object Segmentation from Rigid Dynamics of Point CloudsZiyang Song, Bo YangNeurIPS 2022 · 41 citations
- Self-Supervised Pretraining for Large-Scale Point CloudsZaiwei Zhang, Min Bai, Li Erran LiNeurIPS 2022 · 12 citations
- PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian SplattingYixiao Song, Qingyong Li, Wen Wang, Zhicheng YanCVPR 2026 · 4 citations
- unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary ReasoningYafei Yang, Zihui Zhang, Bo YangICML 2025
- Self-Supervised Learning on 3D Point Clouds by Learning Discrete Generative ModelsBenjamin Eckart, Wentao Yuan, Chao Liu, Jan KautzCVPR 2021
