Learning Object-Centric Representations of Multi-Object Scenes from Multiple Views
Nanbo Li, Cian Eastwood, Robert B. Fisher
Abstract
Learning object-centric representations of multi-object scenes is a promising approach towards machine intelligence, facilitating high-level reasoning and control from visual sensory data. However, current approaches for unsupervised objectcentric scene representation are incapable of aggregating information from multiple observations of a scene. As a result, these "single-view" methods form their representations of a 3D scene based only on a single 2D observation (view). Naturally, this leads to several inaccuracies, with these methods falling victim to single-view spatial ambiguities. To address this, we propose The Multi-View and Multi-Object Network (MulMON)-a method for learning accurate, object-centric representations of multi-object scenes by leveraging multiple views. In order to sidestep the main technical difficulty of the multi-object-multi-view scenario-maintaining object correspondences across views-MulMON iteratively updates the latent object representations for a scene over multiple views. To ensure that these iterative updates do indeed aggregate spatial information to form a complete 3D scene understanding, MulMON is asked to predict the appearance of the scene from novel viewpoints during training. Through experiments we show that MulMON better-resolves spatial ambiguities than single-view methods-learning more accurate and disentangled object representations-and also achieves new functionality in predicting object segmentations for novel viewpoints. Our implementation and pretrained models are given on GitHub 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fa2e8c1-c2a7-4ba7-ae67-a07d3f5d0397Cited by top-tier papers22
- Neural Systematic BinderGautam Singh, Yeongbin Kim, Sungjin AhnICLR 2023 · 105 citations
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video DecompositionRishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey et al.NeurIPS 2021 · 90 citations
- SemanticStyleGAN: Learning Compositional Generative Priors for Controllable Image Synthesis and EditingYichun Shi, Xiao Yang, Yangyue Wan, Xiaohui ShenCVPR 2022 · 88 citations
- Efficient Iterative Amortized Inference for Learning Symmetric and Disentangled Multi-Object RepresentationsPatrick Emami, Pan He, Sanjay Ranka, Anand RangarajanICML 2021 · 48 citations
- Compositional Transformers for Scene GenerationDrew A. Hudson, Larry ZitnickNeurIPS 2021 · 36 citations
Builds on2
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun et al.ICLR 2020 · 276 citations
Related papers
- Object-Centric Representation Learning with Generative Spatial-Temporal FactorizationNanbo Li, Muhammad Ahmed Raza, Wenbin Hu, Zhaole Sun et al.NeurIPS 2021 · 17 citations
- Generalization and Robustness Implications in Object-Centric LearningAndrea Dittadi, Samuele S. Papa, Michele De Vita, Bernhard Schölkopf et al.ICML 2022 · 87 citations
- Identifiable Object Representations under Spatial AmbiguitiesAvinash Kori, Francesca Toni, Ben GlockerICML 2025
- Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified ViewpointsJinyang Yuan, Bin Li, Xiangyang XueAAAI 2022 · 12 citations
- Learning Relationships for Multi-View 3D Object RecognitionZe Yang, Liwei WangICCV 2019 · 166 citations
