Learning Object-Centric Representations of Multi-Object Scenes from Multiple Views
Nanbo Li, Cian Eastwood, Robert B. Fisher
摘要
Learning object-centric representations of multi-object scenes is a promising approach towards machine intelligence, facilitating high-level reasoning and control from visual sensory data. However, current approaches for unsupervised objectcentric scene representation are incapable of aggregating information from multiple observations of a scene. As a result, these "single-view" methods form their representations of a 3D scene based only on a single 2D observation (view). Naturally, this leads to several inaccuracies, with these methods falling victim to single-view spatial ambiguities. To address this, we propose The Multi-View and Multi-Object Network (MulMON)-a method for learning accurate, object-centric representations of multi-object scenes by leveraging multiple views. In order to sidestep the main technical difficulty of the multi-object-multi-view scenario-maintaining object correspondences across views-MulMON iteratively updates the latent object representations for a scene over multiple views. To ensure that these iterative updates do indeed aggregate spatial information to form a complete 3D scene understanding, MulMON is asked to predict the appearance of the scene from novel viewpoints during training. Through experiments we show that MulMON better-resolves spatial ambiguities than single-view methods-learning more accurate and disentangled object representations-and also achieves new functionality in predicting object segmentations for novel viewpoints. Our implementation and pretrained models are given on GitHub 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Neural Systematic BinderGautam Singh, Yeongbin Kim, Sungjin AhnICLR 2023 · 被引用 105 次
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video DecompositionRishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey 等NeurIPS 2021 · 被引用 90 次
- SemanticStyleGAN: Learning Compositional Generative Priors for Controllable Image Synthesis and EditingYichun Shi, Xiao Yang, Yangyue Wan, Xiaohui ShenCVPR 2022 · 被引用 88 次
- Efficient Iterative Amortized Inference for Learning Symmetric and Disentangled Multi-Object RepresentationsPatrick Emami, Pan He, Sanjay Ranka, Anand RangarajanICML 2021 · 被引用 48 次
- Compositional Transformers for Scene GenerationDrew A. Hudson, Larry ZitnickNeurIPS 2021 · 被引用 36 次
它引用的顶会 Paper2
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun 等ICLR 2020 · 被引用 276 次
相关 Paper
- Object-Centric Representation Learning with Generative Spatial-Temporal FactorizationNanbo Li, Muhammad Ahmed Raza, Wenbin Hu, Zhaole Sun 等NeurIPS 2021 · 被引用 17 次
- Generalization and Robustness Implications in Object-Centric LearningAndrea Dittadi, Samuele S. Papa, Michele De Vita, Bernhard Schölkopf 等ICML 2022 · 被引用 87 次
- Identifiable Object Representations under Spatial AmbiguitiesAvinash Kori, Francesca Toni, Ben GlockerICML 2025
- Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified ViewpointsJinyang Yuan, Bin Li, Xiangyang XueAAAI 2022 · 被引用 12 次
- Learning Relationships for Multi-View 3D Object RecognitionZe Yang, Liwei WangICCV 2019 · 被引用 166 次
