DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
Qihao Liu, Yi Zhang, Song Bai, Adam Kortylewski, Alan L. Yuille
摘要
We present DIRECT-3D a diffusion-based 3D generative model for creating high-quality 3D assets (represented by Neural Radiance Fields) from text prompts. Unlike recent 3D generative models that rely on clean and well-aligned 3D data limiting them to single or few-class generation our model is directly trained on extensive noisy and unaligned `in-the-wild' 3D assets mitigating the key challenge (i.e. data scarcity) in large-scale 3D generation. In particular DIRECT-3D is a tri-plane diffusion model that integrates two innovations: 1) A novel learning framework where noisy data are filtered and aligned automatically during the training process. Specifically after an initial warm-up phase using a small set of clean data an iterative optimization is introduced in the diffusion process to explicitly estimate the 3D pose of objects and select beneficial data based on conditional density. 2) An efficient 3D representation that is achieved by disentangling object geometry and color features with two separate conditional diffusion models that are optimized hierarchically. Given a prompt input our model generates high-quality high-resolution realistic and complex 3D objects with accurate geometric details in seconds. We achieve state-of-the-art performance in both single-class generation and text-to-3D generation. We also demonstrate that DIRECT-3D can serve as a useful 3D geometric prior of objects for example to alleviate the well-known Janus problem in 2D-lifting methods such as DreamFusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- TRIM: Scalable 3D Gaussian Diffusion Inference with Temporal and Spatial TrimmingZeyuan Yin, Xiaoming LiuNeurIPS 2025 · 被引用 4 次
- Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D GenerationZiying Li, Xuequan Lu, Xinkui Zhao, Guanjie Cheng 等NeurIPS 2025 · 被引用 4 次
- GaussianGrow: Geometry-aware Gaussian Growing from 3D Point Clouds with Text GuidanceWeiqi Zhang, Junsheng Zhou, Haotian Geng, Kanle Shi 等CVPR 2026 · 被引用 2 次
- Representing 3D Shapes with 64 Latent Vectors for 3D Diffusion ModelsIn Cho, Youngbeom Yoo, Subin Jeon, Seon Joo KimICCV 2025 · 被引用 1 次
- Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality EvolutionQihao Liu, Xi Yin, Alan L. Yuille, Andrew Brown 等CVPR 2025
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- A Unified Approach for Text-and Image-Guided 4D Scene GenerationYufeng Zheng, Xueting Li, Koki Nagano, Sifei Liu 等CVPR 2024 · 被引用 20 次
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang 等NeurIPS 2024 · 被引用 251 次
- DreamBooth3D: Subject-Driven Text-to-3D GenerationAmit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer 等ICCV 2023 · 被引用 280 次
- GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion ModelsTaoran Yi, Jiemin Fang, Junjie Wang, Guanjun Wu 等CVPR 2024 · 被引用 106 次
- Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction ModelJiahao Li, Hao Tan, Kai Zhang, Zexiang Xu 等ICLR 2024 · 被引用 408 次
