Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images
Cuican Yu, Guansong Lu, Yihan Zeng, Jian Sun, Xiaodan Liang, Huibin Li, Zongben Xu, Songcen Xu, Wei Zhang, Hang Xu
摘要
Generating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie, and robotics. Recent progresses have demonstrated the success of unconditional 3D face generation and text-to-3D shape generation. However, due to the limited text-3D face data pairs, text-driven 3D face generation remains an open problem. In this paper, we propose a text-guided 3D faces generation method, refer as TG-3DFace, for generating realistic 3D faces using text guidance. Specifically, we adopt an unconditional 3D face generation framework and equip it with text conditions, which learns the text-guided 3D face generation with only text-2D face data. On top of that, we propose two text-to-face cross-modal alignment techniques, including the global contrastive learning and the fine-grained alignment module, to facilitate high semantic consistency between generated 3D faces and input texts. Besides, we present directional classifier guidance during the inference process, which encourages creativity for out-of-domain generations. Compared to the existing methods, TG-3DFace creates more realistic and aesthetically pleasing 3D faces, boosting 9% multi-view consistency (MVIC) over Latent3D. The rendered face images generated by TG-3DFace achieve higher FID and CLIP score than text-to-2D face/image generation models, demonstrating our superiority in generating realistic and semantic-consistent textures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion PriorsTianyu Huang, Haoze Zhang, Yihan Zeng, Zhilu Zhang 等AAAI 2025 · 被引用 21 次
- HeadArtist: Text-conditioned 3D Head Generation with Self Score DistillationHongyu Liu, Xuan Wang, Ziyu Wan, Yujun Shen 等SIGGRAPH 2024 · 被引用 9 次
- Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric RegularizationJinlu Zhang, Yiyi Zhou, Qiancheng Zheng, Xiaoxiong Du 等ICML 2024 · 被引用 9 次
- InstructAvatar: Text-Guided Emotion and Motion Control for Avatar GenerationYuchi Wang, Junliang Guo, Jianhong Bai, Runyi Yu 等AAAI 2025 · 被引用 5 次
- DreamControl: Control-Based Text-to-3D Generation with 3D Self-PriorTianyu Huang, Yihan Zeng, Zhilu Zhang, Wan Xu 等CVPR 2024
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- Text-Conditional Attribute Alignment Across Latent Spaces for 3D Controllable Face Image SynthesisFeifan Xu, Rui Li, Si Wu, Yong Xu 等CVPR 2024
- Text-Guided 3D Face Synthesis - From Generation to EditingYunjie Wu, Yapeng Meng, Zhipeng Hu, Lincheng Li 等CVPR 2024 · 被引用 11 次
- TANGO: Text-driven Photorealistic and Robust 3D Stylization via Lighting DecompositionYongwei Chen, Rui Chen, Jiabao Lei, Yabin Zhang 等NeurIPS 2022 · 被引用 112 次
- Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationZibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng 等NeurIPS 2023 · 被引用 279 次
- Controllable 3D Face Generation with Conditional Style Code DiffusionXiaolong Shen, Jianxin Ma, Chang Zhou, Zongxin YangAAAI 2024 · 被引用 19 次
