Dreamweaver: Learning Compositional World Models from Pixels
Junyeob Baek, Yi-Fu Wu, Gautam Singh, Sungjin Ahn
Abstract
Humans have an innate ability to decompose their perceptions of the world into objects and their attributes, such as colors, shapes, and movement patterns. This cognitive process enables us to imagine novel futures by recombining familiar concepts. However, replicating this ability in artificial intelligence systems has proven challenging, particularly when it comes to modeling videos into compositional concepts and generating unseen, recomposed futures without relying on auxiliary data, such as text, masks, or bounding boxes. In this paper, we propose Dreamweaver, a neural architecture designed to discover hierarchical and compositional representations from raw videos and generate compositional future simulations. Our approach leverages a novel Recurrent Block-Slot Unit (RBSU) to decompose videos into their constituent objects and attributes. In addition, Dreamweaver uses a multi-future-frame prediction objective to capture disentangled representations for dynamic concepts more effectively as well as static concepts. In experiments, we demonstrate our model outperforms current state-of-the-art baselines for world modeling when evaluated under the DCI framework across multiple datasets. Furthermore, we show how the modularized concept representations of our model enable compositional imagination, allowing the generation of novel videos by recombining attributes from previously seen objects. cun-bjy.github.io/dreamweaver-website
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 907d459c-a986-4e89-bc3f-b8253645ac03Cited by top-tier papers3
- Dyn-O: Building Structured World Models with Object-Centric RepresentationsZizhao Wang, Kaixin Wang, Li Zhao, Peter Stone et al.NeurIPS 2025 · 15 citations
- Learning Interactive World Model for Object-Centric Reinforcement LearningFan Feng, Phillip Lippe, Sara MagliacaneNeurIPS 2025 · 13 citations
- VideoWorld 2: Learning Transferable Knowledge from Real-world VideosZhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo et al.CVPR 2026 · 9 citations
Builds on31
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
Related papers
- RoboDreamer: Learning Compositional World Models for Robot ImaginationSiyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li et al.ICML 2024 · 140 citations
- PARTS: Unsupervised segmentation with slots, attention and independence maximizationDaniel Zoran, Rishabh Kabra, Alexander Lerchner, Danilo J. RezendeICCV 2021 · 53 citations
- Compositional Video Understanding with Spatiotemporal Structure-based TransformersHoyeoung Yun, Jinwoo Ahn, Minseo Kim, Eun-Sol KimCVPR 2024 · 4 citations
- EVOKE: Efficient and High-Fidelity EEG-to-Video Reconstruction via Decoupling Implicit Neural RepresentationHaodong Jing, Panqi Yang, Dongyao Jiang, Zhipeng Liu et al.AAAI 2026 · 1 citation
- Neural Systematic BinderGautam Singh, Yeongbin Kim, Sungjin AhnICLR 2023 · 105 citations
