High-fidelity 3D Face Generation from Natural Language Descriptions
Menghua Wu, Hao Zhu, Linjia Huang, Yiyu Zhuang, Yuanxun Lu, Xun Cao
Abstract
Synthesizing high-quality 3D face models from natural language descriptions is very valuable for many applications, including avatar creation, virtual reality, and telepresence. However, little research ever tapped into this task. We argue the major obstacle lies in 1) the lack of highquality 3D face data with descriptive text annotation, and 2) the complex mapping relationship between descriptive language space and shape/appearance space. To solve these problems, we build DESCRIBE3D dataset, the first large-scale dataset with fine-grained text descriptions for text-to-3D face generation task. Then we propose a twostage framework to first generate a 3D face that matches the concrete descriptions, then optimize the parameters in the 3D shape and texture space with abstract description to refine the 3D face model. Extensive experimental results show that our method can produce a faithful 3D face that conforms to the input descriptions with higher accuracy and quality than previous methods. The code and DE-SCRIBE3D dataset are released at https://github . com/zhuhao-nju/describe3d.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b5df120-2221-49cc-abf9-0ab5db527892Cited by top-tier papers10
- Controllable 3D Face Generation with Conditional Style Code DiffusionXiaolong Shen, Jianxin Ma, Chang Zhou, Zongxin YangAAAI 2024 · 19 citations
- Text-Guided 3D Face Synthesis - From Generation to EditingYunjie Wu, Yapeng Meng, Zhipeng Hu, Lincheng Li et al.CVPR 2024 · 11 citations
- Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric RegularizationJinlu Zhang, Yiyi Zhou, Qiancheng Zheng, Xiaoxiong Du et al.ICML 2024 · 9 citations
- A Pre-convolved Representation for Plug-and-Play Neural Illumination FieldsYiyu Zhuang, Qi Zhang, Xuan Wang, Hao Zhu et al.AAAI 2024 · 3 citations
- StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric PriorsXiaokun Sun, Zeyu Cai, Ying Tai, Jian Yang et al.ICCV 2025 · 3 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
Related papers
- Tera: Rethinking Text-Guided Realistic 3D Avatar GenerationYanwen Wang, Yiyu Zhuang, Jiawei Zhang, Li Wang et al.ICCV 2025 · 2 citations
- Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only ImagesCuican Yu, Guansong Lu, Yihan Zeng, Jian Sun et al.ICCV 2023 · 20 citations
- Towards High-fidelity 3D Talking Avatar with Personalized Dynamic TextureXuanchen Li, Jianyu Wang, Yuhao Cheng, Yikun Zeng et al.CVPR 2025
- TextGaze: Gaze-Controllable Face Generation with Natural LanguageHengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin ChangACM MM 2024 · 3 citations
- DreamHuman: Animatable 3D Avatars from TextNikos Kolotouros, Thiemo Alldieck, Andrei Zanfir, Eduard Gabriel Bazavan et al.NeurIPS 2023 · 136 citations
