RELATE: Physically Plausible Multi-Object Scene Synthesis Using Structured Latent Spaces
Sébastien Ehrhardt, Oliver Groth, Áron Monszpart, Martin Engelcke, Ingmar Posner, Niloy J. Mitra, Andrea Vedaldi
Abstract
We present RELATE, a model that learns to generate physically plausible scenes and videos of multiple interacting objects. Similar to other generative approaches, RELATE is trained end-to-end on raw, unlabeled data. RELATE combines an object-centric GAN formulation with a model that explicitly accounts for correlations between individual objects. This allows the model to generate realistic scenes and videos from a physically-interpretable parameterization. Furthermore, we show that modeling the object correlation is necessary to learn to disentangle object positions and identity. We find that RELATE is also amenable to physically realistic scene editing and that it significantly outperforms prior art in object-centric scene generation in both synthetic (CLEVR, ShapeStacks) and real-world data (cars). In addition, in contrast to state-of-the-art methods in object-centric generative modeling, RELATE also extends naturally to dynamic scenes and generates videos of high visual fidelity. Source code, datasets and more results are available at http://geometry.cs.ucl.ac.uk/projects/2020/relate/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Generative Adversarial TransformersDrew A. Hudson, Larry ZitnickICML 2021 · 213 citations
- Illiterate DALL-E Learns to ComposeGautam Singh, Fei Deng, Sungjin AhnICLR 2022 · 182 citations
- GENESIS-V2: Inferring Unordered Object Representations without Iterative RefinementMartin Engelcke, Oiwi Parker Jones, Ingmar PosnerNeurIPS 2021 · 143 citations
- DiffScene: Diffusion-Based Safety-Critical Scenario Generation for Autonomous VehiclesChejian Xu, Aleksandr Petiushko, Ding Zhao, Bo LiAAAI 2025 · 90 citations
- SemanticStyleGAN: Learning Compositional Generative Priors for Controllable Image Synthesis and EditingYichun Shi, Xiao Yang, Yangyue Wan, Xiaohui ShenCVPR 2022 · 88 citations
Builds on4
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 334 citations
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 84 citations
- Structured Object-Aware Physics Prediction for Video Modeling and PlanningJannik Kossen, Karl Stelzner, Marcel Hussing, Claas Voelcker et al.ICLR 2020 · 77 citations
- Towards Unsupervised Learning of Generative Models for 3D Controllable Image SynthesisYiyi Liao, Katja Schwarz, Lars M. Mescheder, Andreas GeigerCVPR 2020
Related papers
- BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled ImagesThu Nguyen-Phuoc, Christian Richardt, Long Mai, Yong-Liang Yang et al.NeurIPS 2020 · 256 citations
- Interaction-Based Disentanglement of Entities for Object-Centric World ModelsAkihiro Nakano, Masahiro Suzuki, Yutaka MatsuoICLR 2023
- Generative Video Transformer: Can Objects be the Words?Yi-Fu Wu, Jaesik Yoon, Sungjin AhnICML 2021 · 37 citations
- Generative Scene Graph NetworksFei Deng, Zhuo Zhi, Donghun Lee, Sungjin AhnICLR 2021 · 10 citations
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 75 citations
