Beyond Pixel Histories: World Models with Persistent 3D State
Samuel Garcin, Tom Walker, Steven McDonagh, Tim Pearce, Hakan Bilen, Tianyu He, Kaixin Wang, Jiang Bian
摘要
Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, existing models typically lack a 3D representation of the environment, meaning 3D consistency must be implicitly learned from data, and spatial memory is restricted to limited temporal context windows. This results in an unrealistic user experience and presents significant obstacles to downstream tasks such as training agents. To address this, we present PERSIST, a new paradigm of world model which simulates the evolution of a latent 3D scene: environment, camera, and renderer. This allows us to synthesise new frames with persistent spatial memory and consistent geometry. Both quantitative metrics and a qualitative user study show substantial improvements in spatial memory, 3D consistency, and long-horizon stability over existing methods, enabling coherent, evolving 3D worlds. We further demonstrate novel capabilities, including synthesising diverse 3D environments from a single image, as well as enabling fine-grained, geometry-aware control over generated experiences by supporting environment editing and specification directly in 3D space.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz 等NeurIPS 2024 · 被引用 751 次
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 被引用 659 次
相关 Paper
- Video World Models with Long-term Spatial MemoryTong Wu, Shuai Yang, Ryan Po, Yinghao Xu 等NeurIPS 2025 · 被引用 145 次
- Spatia: Video Generation with Updatable Spatial MemoryJinjing Zhao, Fangyun Wei, Zhening Liu, Hongyang Zhang 等CVPR 2026 · 被引用 37 次
- Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-MotionNils Morbitzer, Jonathan Evers, Artem Savkin, Thomas Stauner 等ICML 2026
- WorldMem: Long-term Consistent World Simulation with MemoryZeqi Xiao, Yushi Lan, Yifan Zhou, Wenqi Ouyang 等NeurIPS 2025 · 被引用 165 次
- Learning 3D Persistent Embodied World ModelsSiyuan Zhou, Yilun Du, Yuncong Yang, Lei Han 等NeurIPS 2025 · 被引用 34 次
