DreamFlow: High-quality text-to-3D generation by Approximating Probability Flow
Kyungmin Lee, Kihyuk Sohn, Jinwoo Shin
摘要
Recent progress in text-to-3D generation has been achieved through the utilization of score distillation methods: they make use of the pre-trained text-to-image (T2I) diffusion models by distilling via the diffusion model training objective. However, such an approach inevitably results in the use of random timesteps at each update, which increases the variance of the gradient and ultimately prolongs the optimization process. In this paper, we propose to enhance the text-to-3D optimization by leveraging the T2I diffusion prior in the generative sampling process with a predetermined timestep schedule. To this end, we interpret text-to-3D optimization as a multi-view image-to-image translation problem, and propose a solution by approximating the probability flow. By leveraging the proposed novel optimization algorithm, we design DreamFlow, a practical three-stage coarseto-fine text-to-3D optimization framework that enables fast generation of highquality and high-resolution (i.e., 1024×1024) 3D contents. For example, we demonstrate that DreamFlow is 5 times faster than the existing state-of-the-art text-to-3D method, while producing more photorealistic 3D contents. 1 INTRODUCTION High-quality 3D content generation is crucial for a broad range of applications, including entertainment, gaming, augmented/virtual/mixed reality, and robotics simulation. However, the current 3D generation process entails tedious work with 3D modeling software, which demands a lot of time and expertise. Thereby, 3D generative models (Gao et al., 2022; Chan et al., 2022; Zeng et al., 2022) have brought large attention, yet they are limited by their generalization capability to creative and artistic 3D contents due to the scarcity of high-quality 3D dataset. Recent works have demonstrated the great promise of text-to-3D generation, which enables creative and diverse 3D content creation with textual descriptions (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- One-Step Effective Diffusion Network for Real-World Image Super-ResolutionRongyuan Wu, Lingchen Sun, Zhiyuan Ma, Lei ZhangNeurIPS 2024 · 被引用 319 次
- Rethinking Score Distillation as a Bridge Between Image DistributionsDavid McAllister, Songwei Ge, Jia-Bin Huang, David Jacobs 等NeurIPS 2024 · 被引用 43 次
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang 等ICLR 2026 · 被引用 33 次
- Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D GenerationZiying Li, Xuequan Lu, Xinkui Zhao, Guanjie Cheng 等NeurIPS 2025 · 被引用 4 次
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture InfillingShuhong Zheng, Ashkan Mirzaei, Igor GilitschenskiNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- PI3D: Efficient Text-to-3D Generation with Pseudo-Image DiffusionYing-Tian Liu, Yuan-Chen Guo, Guan Luo, Heyi Sun 等CVPR 2024
- DreamPropeller: Supercharge Text-to-3D Generation with Parallel SamplingLinqi Zhou, Andy Shih, Chenlin Meng, Stefano ErmonCVPR 2024 · 被引用 9 次
- DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D GenerationYukun Huang, Jianan Wang, Yukai Shi, Boshi Tang 等ICLR 2024 · 被引用 79 次
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 被引用 463 次
- Magic3D: High-Resolution Text-to-3D Content CreationChen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa 等CVPR 2023
