Control3D: Towards Controllable Text-to-3D Generation
Yang Chen, Yingwei Pan, Yehao Li, Ting Yao, Tao Mei
Abstract
Recent remarkable advances in large-scale text-to-image diffusion models have inspired a significant breakthrough in text-to-3D generation, pursuing 3D content creation solely from a given text prompt. However, existing text-to-3D techniques lack a crucial ability in the creative process: interactively control and shape the synthetic 3D contents according to users' desired specifications (e.g., sketch). To alleviate this issue, we present the first attempt for text-to-3D generation conditioning on the additional hand-drawn sketch, namely Control3D, which enhances controllability for users. In particular, a 2D conditioned diffusion model (ControlNet) is remoulded to guide the learning of 3D scene parameterized as NeRF, encouraging each view of 3D scene aligned with the given text prompt and hand-drawn sketch. Moreover, we exploit a pre-trained differentiable photo-to-sketch model to directly estimate the sketch of the rendered image over synthetic 3D scene. Such estimated sketch along with each sampled view is further enforced to be geometrically consistent with the given sketch, pursuing better controllable text-to-3D generation. Through extensive experiments, we demonstrate that our proposal can generate accurate and faithful 3D scenes that align closely with the input text prompts and sketches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43495154-72b9-4279-9672-41e3e109abdaCited by top-tier papers19
- Boosting Diffusion Models with Moving Average Sampling in Frequency DomainYurui Qian, Qi Cai, Yingwei Pan, Yehao Li et al.CVPR 2024 · 22 citations
- Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion ModelsHaibo Yang, Yang Chen, Yingwei Pan, Ting Yao et al.ACM MM 2024 · 22 citations
- SD-DiT: Unleashing the Power of Self-Supervised Discrimination in Diffusion Transformer*Rui Zhu, Yingwei Pan, Yehao Li, Ting Yao et al.CVPR 2024 · 15 citations
- Tetrahedron Splatting for 3D GenerationChun Gu, Zeyu Yang, Zijie Pan, Xiatian Zhu et al.NeurIPS 2024 · 13 citations
- X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar GenerationYiwei Ma, Zhekai Lin, Jiayi Ji, Yijun Fan et al.ICML 2024 · 9 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- SketchDream: Sketch-based Text-To-3D Generation and EditingFeng-Lin Liu, Hongbo Fu, Yu-Kun Lai, Lin GaoSIGGRAPH 2024 · 30 citations
- Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D GenerationChaohui Yu, Qiang Zhou, Jingliang Li, Zhe Zhang et al.ACM MM 2023 · 26 citations
- SKED: Sketch-guided Text-based 3D EditingAryan Mikaeili, Or Perel, Mehdi Safaee, Daniel Cohen-Or et al.ICCV 2023 · 83 citations
- Diff3DS: Generating View-Consistent 3D Sketch via Differentiable Curve RenderingYibo Zhang, Lihong Wang, Changqing Zou, Tieru Wu et al.ICLR 2025
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
