Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified Viewpoints
Jinyang Yuan, Bin Li, Xiangyang Xue
Abstract
Visual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may vary greatly with the change of viewpoints. When observing a visual scene that contains multiple objects from multiple viewpoints, humans are able to perceive the scene in a compositional way from each viewpoint, while achieving the so-called ``object constancy'' across different viewpoints, even though the exact viewpoints are untold. This ability is essential for humans to identify the same object while moving and to learn from vision efficiently. It is intriguing to design models that have the similar ability. In this paper, we consider a novel problem of learning compositional scene representations from multiple unspecified viewpoints without using any supervision, and propose a deep generative model which separates latent representations into a viewpoint-independent part and a viewpoint-dependent part to solve this problem. To infer latent representations, the information contained in different viewpoints is iteratively integrated by neural networks. Experiments on several specifically designed synthetic datasets have shown that the proposed method is able to effectively learn from multiple unspecified viewpoints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 736d535b-b765-4945-8081-9960bb8cc2f4Cited by top-tier papers3
- Towards Generative Abstract Reasoning: Completing Raven's Progressive Matrix via Rule Abstraction and SelectionFan Shi, Bin Li, Xiangyang XueICLR 2024 · 5 citations
- Improving Viewpoint-Independent Object-Centric Representations through Active Viewpoint SelectionYinxuan Huang, Chengmin Gao, Bin Li, Xiangyang XueNeurIPS 2024 · 2 citations
- Compositional Law Parsing with Latent Random FunctionsFan Shi, Bin Li, Xiangyang XueICLR 2023
Builds on9
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 334 citations
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun et al.ICLR 2020 · 276 citations
- SCALOR: Generative World Models with Scalable Object RepresentationsJindong Jiang, Sepehr Janghorbani, Gerard de Melo, Sungjin AhnICLR 2020 · 152 citations
- Generative Neurosymbolic MachinesJindong Jiang, Sungjin AhnNeurIPS 2020 · 73 citations
Related papers
- GIRAFFE: Representing Scenes As Compositional Generative Neural Feature FieldsMichael Niemeyer, Andreas GeigerCVPR 2021
- Generative Scene Graph NetworksFei Deng, Zhuo Zhi, Donghun Lee, Sungjin AhnICLR 2021 · 10 citations
- BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled ImagesThu Nguyen-Phuoc, Christian Richardt, Long Mai, Yong-Liang Yang et al.NeurIPS 2020 · 256 citations
- Knowledge-Guided Object Discovery with Acquired Deep ImpressionsJinyang Yuan, Bin Li, Xiangyang XueAAAI 2021 · 6 citations
- Unsupervised object-centric video generation and decomposition in 3DPaul Henderson, Christoph H. LampertNeurIPS 2020 · 41 citations
