InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction From Cluttered Scenes
Zesong Yang, Bangbang Yang, Liyuan Cui, Yuewen Ma, Wenqi Dong, Zhaopeng Cui, Chenxuan Cao, Hujun Bao
Abstract
Humans can naturally identify and mentally complete occluded objects in cluttered environments. However, imparting similar cognitive ability to robotics remains challenging even with advanced reconstruction techniques, which models scenes as undifferentiated wholes and fails to recognize complete object from partial observations. In this paper, we propose InstaScene, a new paradigm towards holistic 3D perception of complex scenes with a primary goal: decomposing arbitrary instances while ensuring complete reconstruction. To achieve precise decomposition, we develop a novel spatial contrastive learning by tracing rasterization of each instance across views, significantly enhancing semantic supervision in cluttered scenes. To overcome incompleteness from limited observations, we introduce in-situ generation that harnesses valuable observations and geometric cues, effectively guiding 3D generative models to reconstruct complete instances that seamlessly align with the real world. Experiments on scene decomposition and object completion across complex real-world and synthetic scenes demonstrate that our method achieves superior decomposition accuracy while producing geometrically faithful and visually intact objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn et al.CVPR 2026 · 24 citations
- ShapeR: Robust Conditional 3D Shape Generation from Casual CapturesYawar Siddiqui, Duncan P. Frost, Samir Aroudj, Armen Avetisyan et al.CVPR 2026 · 19 citations
- SimRecon: SimReady Compositional Scene Reconstruction from Real VideosChong Xia, Kai Zhu, Zizhuo Wang, Fangfu Liu et al.CVPR 2026 · 11 citations
- G4Splat: Geometry-Guided Gaussian Splatting with Generative PriorJunfeng Ni, Yixin Chen, Zhifei Yang, Yu Liu et al.ICLR 2026 · 10 citations
- DiffWind: Physics-Informed Differentiable Modeling of Wind-Driven Object DynamicsYuanhang Lei, Boming Zhao, Zesong Yang, Xingxuan Li et al.ICLR 2026 · 3 citations
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
Related papers
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D ReconstructionHao Li, Zhengyu Zou, Fangfu Liu, Xuanyang Zhang et al.ICLR 2026 · 27 citations
- RevealNet: Seeing Behind Objects in RGB-D ScansJi Hou, Angela Dai, Matthias NießnerCVPR 2020
- I-Scene: 3D Instance Models are Implicit Generalizable Spatial LearnersLu Ling, Yunhao Ge, Yichen Sheng, Aniket BeraCVPR 2026 · 7 citations
- Point-based Instance Completion with Scene ConstraintsWesley Khademi, Fuxin LiICLR 2025
- Instance-Aware Contrastive Learning for Occluded Human Mesh ReconstructionMi-Gyeong Gwon, Gi-Mun Um, Won-Sik Cheong, Wonjun KimCVPR 2024
