SIMstack: A Generative Shape and Instance Model for Unordered Object Stacks
Zoe Landgraf, Raluca Scona, Tristan Laidlow, Stephen James, Stefan Leutenegger, Andrew J. Davison
Abstract
By estimating 3D shape and instances from a single view, we can capture information about an environment quickly, without the need for comprehensive scanning and multi-view fusion. Solving this task for composite scenes (such as object stacks) is challenging: occluded areas are not only ambiguous in shape but also in instance segmentation; multiple decompositions could be valid. We observe that physics constrains decomposition as well as shape in occluded regions and hypothesise that a latent space learned from scenes built under physics simulation can serve as a prior to better predict shape and instances in occluded regions. To this end we propose SIMstack, a depth-conditioned Variational Auto-Encoder (VAE), trained on a dataset of objects stacked under physics simulation. We formulate instance segmentation as a centre voting task which allows for class-agnostic detection and doesn’t require setting the maximum number of objects in the scene. At test time, our model can generate 3D shape and instance segmentation from a single depth view, probabilistically sampling proposals for the occluded region from the learned latent space. Our method has practical applications in providing robots some of the ability humans have to make rapid intuitive inferences of partially observed scenes. We demonstrate an application for precise (non-disruptive) object grasping of unknown objects from a single depth view.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5dfff27c-2df7-4a53-bb87-36c116686e20Builds on11
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou et al.ICCV 2019 · 373 citations
- 3D Scene Reconstruction With Multi-Layer Depth and Epipolar TransformersDaeyun Shin, Zhile Ren, Erik B. Sudderth, Charless C. FowlkesICCV 2019 · 67 citations
- X-Section: Cross-Section Prediction for Enhanced RGB-D FusionAndrea Nicastro, Ronald Clark, Stefan LeuteneggerICCV 2019 · 15 citations
- DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere TracingShaohui Liu, Yinda Zhang, Songyou Peng, Boxin Shi et al.CVPR 2020
Related papers
- Variational Amodal Object CompletionHuan Ling, David Acuna, Karsten Kreis, Seung Wook Kim et al.NeurIPS 2020 · 56 citations
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video DecompositionRishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey et al.NeurIPS 2021 · 90 citations
- Counting Stacked ObjectsCorentin Dumery, Noa Etté, Aoxiang Fan, Ren Li et al.ICCV 2025 · 3 citations
- Sharf: Shape-conditioned Radiance Fields from a Single ViewKonstantinos Rematas, Ricardo Martin-Brualla, Vittorio FerrariICML 2021 · 122 citations
- Unsupervised Learning of Probably Symmetric Deformable 3D Objects From Images in the WildShangzhe Wu, Christian Rupprecht, Andrea VedaldiCVPR 2020
