ATT3D: Amortized Text-to-3D Object Synthesis
Jonathan Lorraine, Kevin Xie, Xiaohui Zeng, Chen-Hsuan Lin, Towaki Takikawa, Nicholas Sharp, Tsung-Yi Lin, Ming-Yu Liu, Sanja Fidler, James Lucas
Abstract
NeRF "A monkey sitting in a chair wearing a suit .. party hat" "A pig riding a motorbike wearing a backpack .. top hat" ...etc... 1hr per prompt 1sec per prompt NeRF text NeRF mapping network trained offline ... ... expensive per-prompt optimization Existing Methods ATT3D: Amortized Text-to-3D Requires 1 hour Requires < 1 sec Figure 1 : Our method initially trains one network to output 3D objects consistent with various text prompts. After, when we receive an unseen prompt, we produce an accurate object in < 1 second, with 1 GPU. Existing methods re-train the entire network for every prompt, requiring a long delay for the optimization to complete. Further, we can interpolate between prompts for user-guided asset generation (Fig. 3 ). We include a project webpage with an overview and videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d3931d5-ce8a-41f7-8990-bcfb2c8d6c87Cited by top-tier papers18
- DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content CreationJiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu et al.ICLR 2024 · 955 citations
- GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion ModelsTaoran Yi, Jiemin Fang, Junjie Wang, Guanjun Wu et al.CVPR 2024 · 106 citations
- IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D GenerationLuke Melas-Kyriazi, Iro Laina, Christian Rupprecht, Natalia Neverova et al.ICML 2024 · 92 citations
- Meta 3D AssetGen: Text-to-Mesh Generation with High-Quality Geometry, Texture, and PBR MaterialsYawar Siddiqui, Tom Monnier, Filippos Kokkinos, Mahendra Kariya et al.NeurIPS 2024 · 89 citations
- MeshFormer : High-Quality Mesh Generation with 3D-Guided Reconstruction ModelMinghua Liu, Chong Zeng, Xinyue Wei, Ruoxi Shi et al.NeurIPS 2024 · 73 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction ModelJiahao Li, Hao Tan, Kai Zhang, Zexiang Xu et al.ICLR 2024 · 408 citations
- Instant3dit: Multiview Inpainting for Fast Editing of 3D ObjectsAmir Barda, Matheus Gadelha, Vladimir G. Kim, Noam Aigerman et al.CVPR 2025
- Shap-Editor: Instruction-guided Latent 3D Editing in SecondsMinghao Chen, Junyu Xie, Iro Laina, Andrea VedaldiCVPR 2024
- Text-To-4D Dynamic Scene GenerationUriel Singer, Shelly Sheynin, Adam Polyak, Oron Ashual et al.ICML 2023 · 234 citations
- Magic3D: High-Resolution Text-to-3D Content CreationChen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa et al.CVPR 2023
