GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
Guanxing Lu, Baoxiong Jia, Puhao Li, Yixin Chen, Ziwei Wang, Yansong Tang, Siyuan Huang
Abstract
Training robot policies within a learned world model is trending due to the inefficiency of real-world interactions. The established image-based world models and policies have shown prior success, but lack robust geometric information that requires consistent spatial and physical understanding of the three-dimensional world, even pretrained on internet-scale video sources. To this end, we propose a novel branch of world model named Gaussian World Model (GWM) for robotic manipulation, which reconstructs the future state by inferring the propagation of Gaussian primitives under the effect of robot actions. At its core is a latent Diffusion Transformer (DiT) combined with a 3D variational autoencoder, enabling fine-grained scenelevel future state reconstruction with Gaussian Splatting. GWM can not only enhance the visual representation for imitation learning agent by self-supervised future prediction training, but can serve as a neural simulator that supports model-based reinforcement learning. Both simulated and real-world experiments depict that GWM can precisely predict future scenes conditioned on diverse robot actions, and can be further utilized to train policies that outperform the state-of-the-art by impressive margins, showcasing the initial data scaling potential of 3D world model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ae35b79-6545-4975-b2f2-a4383ce5dc6eCited by top-tier papers7
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic ManipulationWenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu et al.CVPR 2026 · 87 citations
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA ManipulationJingjing Qian, Boyao Han, Chen Shi, Lei Xiao et al.CVPR 2026 · 19 citations
- Learning 3D Representations for Spatial Intelligence from Unposed Multi-View ImagesBo Zhou, Qiuxia Lai, Zeren Sun, Xiangbo Shu et al.CVPR 2026 · 1 citation
- OMP: One-step Meanflow Policy with Directional AlignmentHan Fang, Yize Huang, Yuheng Zhao, Paul Weng et al.ICML 2026
- MSCD-GS: Motion-Separated Cooperative Deblurring Dynamic Reconstruction via Gaussian Splattingyongjian liao, Xu Zou, Wenjun Chen, Huixuan Li et al.CVPR 2026
Builds on37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
Related papers
- GSV3D: Gaussian Splatting-Based Geometric Distillation With Stable Video Diffusion for Single-Image 3D Object GenerationYe Tao, Jiawei Zhang, Yahao Shi, Dongqing Zou et al.ICCV 2025
- Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene GenerationPanwang Pan, Chenguo Lin, Chenxin Li, Jingjing Zhao et al.CVPR 2026
- Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete DiffusionLunjun Zhang, Yuwen Xiong, Ze Yang, Sergio Casas et al.ICLR 2024 · 105 citations
- Spatial-Temporal Aware Visuomotor Diffusion Policy LearningZhenyang Liu, Yikai Wang, Kuanning Wang, Longfei Liang et al.ICCV 2025 · 11 citations
- DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic ManipulationKaizhao Zhang, Tian Niu, Tianyu Liu, Chenen Guo et al.CVPR 2026
