Improving Generative Imagination in Object-Centric World Models
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Bofeng Fu, Jindong Jiang, Sungjin Ahn
Abstract
The remarkable recent advances in object-centric generative world models raise a few questions. First, while many of the recent achievements are indispensable for making a general and versatile world model, it is quite unclear how these ingredients can be integrated into a unified framework. Second, despite using generative objectives, abilities for object detection and tracking are mainly investigated, leaving the crucial ability of temporal imagination largely under question. Third, a few key abilities for more faithful temporal imagination such as multimodal uncertainty and situationawareness are missing. In this paper, we introduce Generative Structured World Models (G-SWM). The G-SWM achieves the versatile world modeling not only by unifying the key properties of previous models in a principled framework but also by achieving two crucial new abilities, multimodal uncertainty and situation-awareness. Our thorough investigation on the temporal generation ability in comparison to the previous models demonstrates that G-SWM achieves the versatility with the best or comparable performance for all experiment settings including a few complex settings that have not been tested before. https: //sites.google.com/view/gswm
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d72d568b-b7b8-40bd-a93c-e510f03547b3Cited by top-tier papers41
- Simple Unsupervised Object-Centric Learning for Complex and Naturalistic VideosGautam Singh, Yi-Fu Wu, Sungjin AhnNeurIPS 2022 · 182 citations
- Illiterate DALL-E Learns to ComposeGautam Singh, Fei Deng, Sungjin AhnICLR 2022 · 182 citations
- Object-Centric Slot DiffusionJindong Jiang, Fei Deng, Gautam Singh, Sungjin AhnNeurIPS 2023 · 106 citations
- SlotDiffusion: Object-Centric Generative Modeling with Diffusion ModelsZiyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski et al.NeurIPS 2023 · 106 citations
- Neural Systematic BinderGautam Singh, Yeongbin Kim, Sungjin AhnICLR 2023 · 105 citations
Builds on4
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun et al.ICLR 2020 · 276 citations
- Structured Object-Aware Physics Prediction for Video Modeling and PlanningJannik Kossen, Karl Stelzner, Marcel Hussing, Claas Voelcker et al.ICLR 2020 · 77 citations
- Exploiting Spatial Invariance for Scalable Unsupervised Object TrackingEric Crawford, Joelle PineauAAAI 2020 · 68 citations
Related papers
- Structured World Belief for Reinforcement Learning in POMDPGautam Singh, Skand Vishwanath Peri, Junghyun Kim, Hyunseok Kim et al.ICML 2021 · 34 citations
- Generative Neurosymbolic MachinesJindong Jiang, Sungjin AhnNeurIPS 2020 · 73 citations
- Dyn-O: Building Structured World Models with Object-Centric RepresentationsZizhao Wang, Kaixin Wang, Li Zhao, Peter Stone et al.NeurIPS 2025 · 15 citations
- DreamTrack: Dreaming the Future for Multimodal Visual Object TrackingMingzhe Guo, Weiping Tan, Wenyu Ran, Liping Jing et al.CVPR 2025
- Slot-VAE: Object-Centric Scene Generation with Slot AttentionYanbo Wang, Letao Liu, Justin DauwelsICML 2023 · 29 citations
