Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified Viewpoints
Jinyang Yuan, Bin Li, Xiangyang Xue
摘要
Visual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may vary greatly with the change of viewpoints. When observing a visual scene that contains multiple objects from multiple viewpoints, humans are able to perceive the scene in a compositional way from each viewpoint, while achieving the so-called ``object constancy'' across different viewpoints, even though the exact viewpoints are untold. This ability is essential for humans to identify the same object while moving and to learn from vision efficiently. It is intriguing to design models that have the similar ability. In this paper, we consider a novel problem of learning compositional scene representations from multiple unspecified viewpoints without using any supervision, and propose a deep generative model which separates latent representations into a viewpoint-independent part and a viewpoint-dependent part to solve this problem. To infer latent representations, the information contained in different viewpoints is iteratively integrated by neural networks. Experiments on several specifically designed synthetic datasets have shown that the proposed method is able to effectively learn from multiple unspecified viewpoints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Generative Abstract Reasoning: Completing Raven's Progressive Matrix via Rule Abstraction and SelectionFan Shi, Bin Li, Xiangyang XueICLR 2024 · 被引用 5 次
- Improving Viewpoint-Independent Object-Centric Representations through Active Viewpoint SelectionYinxuan Huang, Chengmin Gao, Bin Li, Xiangyang XueNeurIPS 2024 · 被引用 2 次
- Compositional Law Parsing with Latent Random FunctionsFan Shi, Bin Li, Xiangyang XueICLR 2023
它引用的顶会 Paper9
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 被引用 334 次
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun 等ICLR 2020 · 被引用 276 次
- SCALOR: Generative World Models with Scalable Object RepresentationsJindong Jiang, Sepehr Janghorbani, Gerard de Melo, Sungjin AhnICLR 2020 · 被引用 152 次
- Generative Neurosymbolic MachinesJindong Jiang, Sungjin AhnNeurIPS 2020 · 被引用 73 次
相关 Paper
- GIRAFFE: Representing Scenes As Compositional Generative Neural Feature FieldsMichael Niemeyer, Andreas GeigerCVPR 2021
- Generative Scene Graph NetworksFei Deng, Zhuo Zhi, Donghun Lee, Sungjin AhnICLR 2021 · 被引用 10 次
- BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled ImagesThu Nguyen-Phuoc, Christian Richardt, Long Mai, Yong-Liang Yang 等NeurIPS 2020 · 被引用 256 次
- Knowledge-Guided Object Discovery with Acquired Deep ImpressionsJinyang Yuan, Bin Li, Xiangyang XueAAAI 2021 · 被引用 6 次
- Unsupervised object-centric video generation and decomposition in 3DPaul Henderson, Christoph H. LampertNeurIPS 2020 · 被引用 41 次
