WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
Ziyue Zhu, Zhanqian Wu, Zhenxin Zhu, Lijun Zhou, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Jin Xie, Jian Yang
Abstract
Recent advances in driving-scene generation and reconstruction have demonstrated significant potential for enhancing autonomous driving systems by producing scalable and controllable training data. Existing generation methods primarily focus on synthesizing diverse and high-fidelity driving videos; however, due to limited 3D consistency and sparse viewpoint coverage, they struggle to support convenient and high-quality novel-view synthesis (NVS). Conversely, recent 3D/4D reconstruction approaches have significantly improved NVS for real-world driving scenes, yet inherently lack generative capabilities. To overcome this dilemma between scene generation and reconstruction, we propose WorldSplat, a novel feed-forward framework for 4D driving-scene generation. Our approach effectively generates consistent multi-track videos through two key steps: (i) We introduce a 4D-aware latent diffusion model integrating multi-modal information to produce pixel-aligned 4D Gaussians in a feed-forward manner. (ii) Subsequently, we refine the novel view videos rendered from these Gaussians using a enhanced video diffusion model. Extensive experiments conducted on benchmark datasets demonstrate that WorldSplat effectively generates high-fidelity, temporally and spatially consistent multi-track novel view driving videos. Project: https://wm-research.github.io/worldsplat/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 229d8697-6d32-483c-8cbe-0bc93a8fecb0Cited by top-tier papers1
Ask how each one uses itBuilds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta et al.NeurIPS 2024 · 403 citations
Related papers
- UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene ReconstructionChen Shi, Shaoshuai Shi, Xiaoyang Lyu, Chunyang Liu et al.ICLR 2026 · 10 citations
- MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse ViewsYuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang et al.NeurIPS 2024 · 126 citations
- Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene GenerationPanwang Pan, Chenguo Lin, Chenxin Li, Jingjing Zhao et al.CVPR 2026
- VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion PriorsJimin Tang, Wenyuan Zhang, Junsheng Zhou, Zian Huang et al.SIGGRAPH 2026
- Glad: A Streaming Scene Generator for Autonomous DrivingBin Xie, Yingfei Liu, Tiancai Wang, Jiale Cao et al.ICLR 2025
