GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
Tianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen, Yuyao Xu, Lijin Yang, Le Xu, Yu Zhang, Bo Zhang, Wuxiong Huang, Hesheng Wang
Abstract
Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to interpret or reason about the driving environment. Moreover, current approaches represent 3D spatial information with point cloud or BEV features do not accurately align textual information with the underlying 3D scene. To address these limitations, we propose a novel unified DWM framework based on 3D Gaussian scene representation, which enables both 3D scene understanding and multi-modal scene generation, while also enabling contextual enrichment for understanding and generation tasks. Our approach directly aligns textual information with the 3D scene by embedding rich linguistic features into each Gaussian primitive, thereby achieving early modality alignment. In addition, we design a novel task-aware language-guided token sampling strategy that removes redundant 3D information and injects accurate and compact 3D tokens into textual understanding.Furthermore, we design a dual-condition multi-modal generation model, where the information captured by our vision-language model is leveraged as a high-level language condition in combination with a low-level image condition, jointly guiding the multi-modal generation process. We conduct comprehensive studies on the nuScenes, OmniDrive-nuScenes, and NuInteract datasets to validate the effectiveness of our framework. Our method achieves state-of-the-art performance. We will release the code publicly on GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99a31a54-bcba-4337-8874-a7f3a46367baCited by top-tier papers6
- Self-Destructive Language ModelsYuhui Wang, Rongyi Zhu, Ting WangICLR 2026 · 14 citations
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal ReasoningYongxin Wang, Zhicheng Yang, Meng Cao, Mingfei Han et al.CVPR 2026
- GeniNav: Generative Model Driven Image-Goal Navigation via Imagination-Guided Consistency Flow MatchingYuqi Chen, Junjie Gao, Yongzhou Pan, Siyuan Song et al.CVPR 2026
- FM-Steer: Enhance Generalist Policies with Value-Guided Cascaded DenoisingHaoming Song, Delin Qu, Yuanqi Yao, Qizhi Chen et al.CVPR 2026
- Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot LearningYinan Deng, Kejia Hu, Ye Chen, Jianyu Dou et al.CVPR 2026
Builds on44
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
Related papers
- HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and GenerationXin Zhou, Dingkang Liang, Sifan Tu, Xiwu Chen et al.ICCV 2025 · 5 citations
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic EncodingYueming Xu, Jiahui Zhang, Ze Huang, Yurui Chen et al.ICLR 2026 · 8 citations
- UniGS: Unified Language-Image-3D Pretraining with Gaussian SplattingHaoyuan Li, Yanpeng Zhou, Tao Tang, Jifei Song et al.ICLR 2025
- TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D AlignmentJiarun Liu, Qifeng Chen, Yiru Zhao, Minghua Liu et al.ICLR 2026 · 1 citation
- DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous DrivingYiyao Zhu, Ying Xue, Haiming Zhang, Guangfeng Jiang et al.CVPR 2026 · 2 citations
