HeadSculpt: Crafting 3D Head Avatars with Text
Xiao Han, Yukang Cao, Kai Han, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song, Tao Xiang, Kwan-Yee K. Wong
Abstract
Recently, text-guided 3D generative methods have made remarkable advancements in producing high-quality textures and geometry, capitalizing on the proliferation of large vision-language and image diffusion models. However, existing methods still struggle to create high-fidelity 3D head avatars in two aspects: (1) They rely mostly on a pre-trained text-to-image diffusion model whilst missing the necessary 3D awareness and head priors. This makes them prone to inconsistency and geometric distortions in the generated avatars. (2) They fall short in fine-grained editing. This is primarily due to the inherited limitations from the pre-trained 2D image diffusion models, which become more pronounced when it comes to 3D head avatars. In this work, we address these challenges by introducing a versatile coarse-to-fine pipeline dubbed HeadSculpt for crafting (i.e., generating and editing) 3D head avatars from textual prompts. Specifically, we first equip the diffusion model with 3D awareness by leveraging landmark-based control and a learned textual embedding representing the back view appearance of heads, enabling 3D-consistent head avatar generations. We further propose a novel identity-aware editing score distillation strategy to optimize a textured mesh with a high-resolution differentiable rendering technique. This enables identity preservation while following the editing instruction. We showcase HeadSculpt's superior fidelity and editing capabilities through comprehensive experiments and comparisons with existing methods. ‡ * Equal contributions † Corresponding authors ‡ Webpage: https://brandonhan.uk/HeadSculpt Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 441d4bf0-4430-49a2-9f53-bac4d9495a40Cited by top-tier papers16
- GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion ModelsTaoran Yi, Jiemin Fang, Junjie Wang, Guanjun Wu et al.CVPR 2024 · 106 citations
- AvatarVerse: High-Quality & Stable 3D Avatar Creation from Text and PoseHuichao Zhang, Bowen Chen, Hao Yang, Liao Qu et al.AAAI 2024 · 73 citations
- Portrait3D: Text-Guided High-Quality 3D Portrait Generation Using Pyramid Representation and GANs PriorYiqian Wu, Hao Xu, Xiangjun Tang, Xien Chen et al.SIGGRAPH 2024 · 14 citations
- HeadArtist: Text-conditioned 3D Head Generation with Self Score DistillationHongyu Liu, Xuan Wang, Ziyu Wan, Yujun Shen et al.SIGGRAPH 2024 · 9 citations
- ID-Sculpt: ID-aware 3D Head Generation from Single In-the-wild Portrait ImageJinkun Hao, Junshu Tang, Jiangning Zhang, Ran Yi et al.AAAI 2025 · 6 citations
Builds on52
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Text-based Animatable 3D Avatars with Morphable Model AlignmentYiqian Wu, Malte Prinzler, Xiaogang Jin, Siyu TangSIGGRAPH 2025 · 1 citation
- ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation SamplingFrancesca Babiloni, Alexandros Lattas, Jiankang Deng, Stefanos ZafeiriouNeurIPS 2024 · 5 citations
- AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose ControlRuixiang Jiang, Can Wang, Jingbo Zhang, Menglei Chai et al.ICCV 2023 · 103 citations
- VP3D: Unleashing 2D Visual Prompt for Text-to-3D GenerationYang Chen, Yingwei Pan, Haibo Yang, Ting Yao et al.CVPR 2024
- StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric PriorsXiaokun Sun, Zeyu Cai, Ying Tai, Jian Yang et al.ICCV 2025 · 3 citations
