Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
Jinlu Zhang, Yiyi Zhou, Qiancheng Zheng, Xiaoxiong Du, Gen Luo, Jun Peng, Xiaoshuai Sun, Rongrong Ji
Abstract
Text-to-3D-aware face (T3D Face) generation and manipulation is an emerging research hot spot in machine learning, which still suffers from low efficiency and poor quality. In this paper, we propose an End-to-End Efficient and Effective network for fast and accurate T3D face generation and manipulation, termed -FaceNet. Different from existing complex generation paradigms, -FaceNet resorts to a direct mapping from text instructions to 3D-aware visual space. We introduce a novel Style Code Enhancer to enhance cross-modal semantic alignment, alongside an innovative Geometric Regularization objective to maintain consistency across multi-view generations. Extensive experiments on three benchmark datasets demonstrate that -FaceNet can not only achieve picture-like 3D face generation and manipulation, but also improve inference speed by orders of magnitudes. For instance, compared with Latent3D, -FaceNet speeds up the five-view generations by almost 470 times, while still exceeding in generation quality. Our code is released at https://github.com/Aria-Zhangjl/E3-FaceNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f69d01b5-e493-4e6a-b4f3-21d28ad876ecCited by top-tier papers3
- CustomTex: High-fidelity Indoor Scene Texturing via Multi-Reference CustomizationWeilin Chen, Jiahao Rao, Wenhao Wang, Xinyang Li et al.CVPR 2026 · 1 citation
- Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial GenerationYushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou et al.AAAI 2026
- ACL: Activating Capability of Linear Attention for Image RestorationYubin Gu, Yuan Meng, Jiayi Ji, Xiaoshuai SunCVPR 2025
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- StyleNeRF: A Style-based 3D Aware Generator for High-resolution Image SynthesisJiatao Gu, Lingjie Liu, Peng Wang, Christian TheobaltICLR 2022 · 622 citations
Related papers
- Towards Open-Ended Text-to-Face Generation, Combination and ManipulationJun Peng, Han Pan, Yiyi Zhou, Jing He et al.ACM MM 2022 · 7 citations
- Text-Conditional Attribute Alignment Across Latent Spaces for 3D Controllable Face Image SynthesisFeifan Xu, Rui Li, Si Wu, Yong Xu et al.CVPR 2024
- Controllable 3D Face Generation with Conditional Style Code DiffusionXiaolong Shen, Jianxin Ma, Chang Zhou, Zongxin YangAAAI 2024 · 19 citations
- Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only ImagesCuican Yu, Guansong Lu, Yihan Zeng, Jian Sun et al.ICCV 2023 · 20 citations
- 3D-Aware Face Editing via Warping-Guided Latent Direction LearningYuhao Cheng, Zhuo Chen, Xingyu Ren, Wenhan Zhu et al.CVPR 2024
