DreamFlow: High-quality text-to-3D generation by Approximating Probability Flow
Kyungmin Lee, Kihyuk Sohn, Jinwoo Shin
Abstract
Recent progress in text-to-3D generation has been achieved through the utilization of score distillation methods: they make use of the pre-trained text-to-image (T2I) diffusion models by distilling via the diffusion model training objective. However, such an approach inevitably results in the use of random timesteps at each update, which increases the variance of the gradient and ultimately prolongs the optimization process. In this paper, we propose to enhance the text-to-3D optimization by leveraging the T2I diffusion prior in the generative sampling process with a predetermined timestep schedule. To this end, we interpret text-to-3D optimization as a multi-view image-to-image translation problem, and propose a solution by approximating the probability flow. By leveraging the proposed novel optimization algorithm, we design DreamFlow, a practical three-stage coarseto-fine text-to-3D optimization framework that enables fast generation of highquality and high-resolution (i.e., 1024×1024) 3D contents. For example, we demonstrate that DreamFlow is 5 times faster than the existing state-of-the-art text-to-3D method, while producing more photorealistic 3D contents. 1 INTRODUCTION High-quality 3D content generation is crucial for a broad range of applications, including entertainment, gaming, augmented/virtual/mixed reality, and robotics simulation. However, the current 3D generation process entails tedious work with 3D modeling software, which demands a lot of time and expertise. Thereby, 3D generative models (Gao et al., 2022; Chan et al., 2022; Zeng et al., 2022) have brought large attention, yet they are limited by their generalization capability to creative and artistic 3D contents due to the scarcity of high-quality 3D dataset. Recent works have demonstrated the great promise of text-to-3D generation, which enables creative and diverse 3D content creation with textual descriptions (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e41069cc-c084-4553-b9fc-d063185b28c1Cited by top-tier papers14
- One-Step Effective Diffusion Network for Real-World Image Super-ResolutionRongyuan Wu, Lingchen Sun, Zhiyuan Ma, Lei ZhangNeurIPS 2024 · 319 citations
- Rethinking Score Distillation as a Bridge Between Image DistributionsDavid McAllister, Songwei Ge, Jia-Bin Huang, David Jacobs et al.NeurIPS 2024 · 43 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
- Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D GenerationZiying Li, Xuequan Lu, Xinkui Zhao, Guanjie Cheng et al.NeurIPS 2025 · 4 citations
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture InfillingShuhong Zheng, Ashkan Mirzaei, Igor GilitschenskiNeurIPS 2025 · 2 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- PI3D: Efficient Text-to-3D Generation with Pseudo-Image DiffusionYing-Tian Liu, Yuan-Chen Guo, Guan Luo, Heyi Sun et al.CVPR 2024
- DreamPropeller: Supercharge Text-to-3D Generation with Parallel SamplingLinqi Zhou, Andy Shih, Chenlin Meng, Stefano ErmonCVPR 2024 · 9 citations
- DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D GenerationYukun Huang, Jianan Wang, Yukai Shi, Boshi Tang et al.ICLR 2024 · 79 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
- Magic3D: High-Resolution Text-to-3D Content CreationChen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa et al.CVPR 2023
