3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
Songchun Zhang, Yibo Zhang, Quan Zheng, Rui Ma, Wei Hua, Hujun Bao, Weiwei Xu, Changqing Zou
Abstract
Text-driven 3D scene generation techniques have made rapid progress in recent years. Their success is mainly at-tributed to using existing generative models to iteratively perform image warping and inpainting to generate 3D scenes. However, these methods heavily rely on the out-puts of existing models, leading to error accumulation in geometry and appearance that prevent the models from being used in various scenarios (e.g., outdoor and unreal sce-narios). To address this limitation, we generatively refine the newly generated local views by querying and aggregating global 3D information, and then progressively generate the 3D scene. Specifically, we employ a tri-plane features-based NeRF as a unified representation of the 3D scene to constrain global 3D consistency, and propose a generative refinement network to synthesize new contents with higher quality by exploiting the natural image prior from 2D dif-fusion model as well as the global 3D information of the current scene. Our extensive experiments demonstrate that, in comparison to previous methods, our approach supports wide variety of scene generation and arbitrary camera tra-jectories with improved visual quality and 3D consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fa572ac-f3c6-407f-b8d3-e460a3d249c1Cited by top-tier papers9
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn et al.CVPR 2026 · 24 citations
- SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited ObservationsSongchun Zhang, Huiyao Xu, Sitong Guo, Zhongwei Xie et al.ICCV 2025 · 6 citations
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry GroundingJulian Ost, Andrea Ramazzina, Amogh Joshi, Maximilian Bömer et al.AAAI 2026 · 6 citations
- Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent DiffusionTongyan Hua, Lutao Jiang, Ying-Cong Chen, Wufan ZhaoICCV 2025 · 5 citations
- WonderTurbo: Generating Interactive 3D World in 0.72 SecondsChaojun Ni, Xiaofeng Wang, Zheng Zhu, Weijie Wang et al.ICCV 2025 · 5 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
Related papers
- SketchDream: Sketch-based Text-To-3D Generation and EditingFeng-Lin Liu, Hongbo Fu, Yu-Kun Lai, Lin GaoSIGGRAPH 2024 · 30 citations
- Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D GenerationChaohui Yu, Qiang Zhou, Jingliang Li, Zhe Zhang et al.ACM MM 2023 · 26 citations
- Real3D: The Curious Case of Neural Scene DegenerationDengsheng Chen, Jie Hu, Xiaoming Wei, Enhua WuAAAI 2024 · 1 citation
- G4Splat: Geometry-Guided Gaussian Splatting with Generative PriorJunfeng Ni, Yixin Chen, Zhifei Yang, Yu Liu et al.ICLR 2026 · 10 citations
- SIGNeRF: Scene Integrated Generation for Neural Radiance FieldsJan-Niklas Dihlmann, Andreas Engelhardt, Hendrik P. A. LenschCVPR 2024
