Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images
Cuican Yu, Guansong Lu, Yihan Zeng, Jian Sun, Xiaodan Liang, Huibin Li, Zongben Xu, Songcen Xu, Wei Zhang, Hang Xu
Abstract
Generating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie, and robotics. Recent progresses have demonstrated the success of unconditional 3D face generation and text-to-3D shape generation. However, due to the limited text-3D face data pairs, text-driven 3D face generation remains an open problem. In this paper, we propose a text-guided 3D faces generation method, refer as TG-3DFace, for generating realistic 3D faces using text guidance. Specifically, we adopt an unconditional 3D face generation framework and equip it with text conditions, which learns the text-guided 3D face generation with only text-2D face data. On top of that, we propose two text-to-face cross-modal alignment techniques, including the global contrastive learning and the fine-grained alignment module, to facilitate high semantic consistency between generated 3D faces and input texts. Besides, we present directional classifier guidance during the inference process, which encourages creativity for out-of-domain generations. Compared to the existing methods, TG-3DFace creates more realistic and aesthetically pleasing 3D faces, boosting 9% multi-view consistency (MVIC) over Latent3D. The rendered face images generated by TG-3DFace achieve higher FID and CLIP score than text-to-2D face/image generation models, demonstrating our superiority in generating realistic and semantic-consistent textures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion PriorsTianyu Huang, Haoze Zhang, Yihan Zeng, Zhilu Zhang et al.AAAI 2025 · 21 citations
- HeadArtist: Text-conditioned 3D Head Generation with Self Score DistillationHongyu Liu, Xuan Wang, Ziyu Wan, Yujun Shen et al.SIGGRAPH 2024 · 9 citations
- Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric RegularizationJinlu Zhang, Yiyi Zhou, Qiancheng Zheng, Xiaoxiong Du et al.ICML 2024 · 9 citations
- InstructAvatar: Text-Guided Emotion and Motion Control for Avatar GenerationYuchi Wang, Junliang Guo, Jianhong Bai, Runyi Yu et al.AAAI 2025 · 5 citations
- DreamControl: Control-Based Text-to-3D Generation with 3D Self-PriorTianyu Huang, Yihan Zeng, Zhilu Zhang, Wan Xu et al.CVPR 2024
Builds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- Text-Conditional Attribute Alignment Across Latent Spaces for 3D Controllable Face Image SynthesisFeifan Xu, Rui Li, Si Wu, Yong Xu et al.CVPR 2024
- Text-Guided 3D Face Synthesis - From Generation to EditingYunjie Wu, Yapeng Meng, Zhipeng Hu, Lincheng Li et al.CVPR 2024 · 11 citations
- TANGO: Text-driven Photorealistic and Robust 3D Stylization via Lighting DecompositionYongwei Chen, Rui Chen, Jiabao Lei, Yabin Zhang et al.NeurIPS 2022 · 112 citations
- Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationZibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng et al.NeurIPS 2023 · 279 citations
- Controllable 3D Face Generation with Conditional Style Code DiffusionXiaolong Shen, Jianxin Ma, Chang Zhou, Zongxin YangAAAI 2024 · 19 citations
