Baking Gaussian Splatting Into Diffusion Denoiser for Fast and Scalable Single-Stage Image-to-3D Generation and Reconstruction
Yuanhao Cai, He Zhang, Kai Zhang, Yixun Liang, Mengwei Ren, Fujun Luan, Qing Liu, Soo Ye Kim, Jianming Zhang, Zhifei Zhang, Yuqian Zhou, Yulun Zhang
Abstract
Existing feedforward image-to-3D methods mainly rely on 2D multi-view diffusion models that cannot guarantee 3D consistency. These methods easily collapse when changing the prompt view direction and mainly handle object-centric cases. In this paper, we propose a novel single-stage 3D diffusion model, DiffusionGS, for object generation and scene reconstruction from a single view. DiffusionGS directly outputs 3D Gaussian point clouds at each timestep to enforce view consistency and allow the model to generate robustly given prompt views of any directions, beyond object-centric inputs. Plus, to improve the capability and generality of DiffusionGS, we scale up 3D training data by developing a scene-object mixed training strategy. Experiments show that DiffusionGS yields improvements of 2.20 dB/23.25 and 1.34 dB/19.16 in PSNR/FID for objects and scenes than the state-of-the-art methods, without depth estimator. Plus, our method enjoys over faster speed ( on an A100 GPU). Project page: https://caiyuanhao1998.github.io/project/DiffusionGS/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 559d8264-c98f-450c-8d3a-6505dd4906faCited by top-tier papers5
- Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D GenerationZiying Li, Xuequan Lu, Xinkui Zhao, Guanjie Cheng et al.NeurIPS 2025 · 4 citations
- S2D: Sparse to Dense Lifting for 3D Reconstruction with Minimal InputsYuzhou Ji, Qijian Tian, He Zhu, Xiaoqi Jiang et al.CVPR 2026 · 1 citation
- Learning Hierarchical Hyperbolic Mixture Model for Part-aware 3D GenerationQitong Yang, Mingtao Feng, Zijie Wu, Huixin Zhu et al.CVPR 2026
- Captain Safari: A World Engine with Pose-Aligned 3D MemoryYu-Cheng Chou, Xingrui Wang, Yitong Li, Jiahao Wang et al.CVPR 2026
- CoverPruneGS: Coverage-Preserving Structured Pruning for Hierarchical 3D Gaussian Splatting from Sparse-View Monocular VideosYang Xiao, Guoan Xu, Guxue Gao, Qiang Wu et al.ICML 2026
Builds on52
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Vistadream: Sampling Multiview Consistent Images for Single-View Scene ReconstructionHaiping Wang, Yuan Liu, Ziwei Liu, Wenping Wang et al.ICCV 2025 · 8 citations
- RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and GenerationTitas Anciukevicius, Zexiang Xu, Matthew Fisher, Paul Henderson et al.CVPR 2023
- Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with TransformersZi-Xin Zou, Zhipeng Yu, Yuan-Chen Guo, Yangguang Li et al.CVPR 2024 · 119 citations
- Generative Gaussian Splatting: Generating 3D Scenes with Video Diffusion PriorsKatja Schwarz, Norman Müller, Peter KontschiederICCV 2025 · 3 citations
- Wonderland: Navigating 3D Scenes from a Single ImageHanwen Liang, Junli Cao, Vidit Goel, Guocheng Qian et al.CVPR 2025
