Text-Conditional Attribute Alignment Across Latent Spaces for 3D Controllable Face Image Synthesis
Feifan Xu, Rui Li, Si Wu, Yong Xu, Hau-San Wong
Abstract
With the advent of generative models and vision-language pre-training, significant improvement has been made in text-driven face manipulation. The text embedding can be used as target supervision for expression control. However, it is non-trivial to associate with its 3D attributes, i.e., pose and illumination. To address these issues, we propose a Text-conditional Attribute aLignment approach for 3D controllable face image synthesis, and our model is referred to as TcALign. Specifically, since the 3D rendered image can be precisely controlled with the 3D face representation, we first propose a Text-conditional 3D Editor to produce the target face representation to realize text-driven manipulation in the 3D space. An attribute embedding space spanned by the target-related attributes embeddings is also introduced to infer the disentangled task-specific direction. Next, we train a cross-modal latent mapping network conditioned on the derived difference of 3D representation to infer a correct vector in the latent space of Style-GAN. This correction vector learning design can accurately transfer the attribute manipulation on 3D images to 2D images. We show that the proposed method delivers more precise text-driven multi-attribute manipulation for 3D controllable face image synthesis. Extensive qualitative and quantitative experiments verify the effectiveness and superiority of our method over the other competing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 528ca39b-164e-4eb0-a159-c42df60cc5e9Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
Related papers
- Towards Open-Ended Text-to-Face Generation, Combination and ManipulationJun Peng, Han Pan, Yiyi Zhou, Jing He et al.ACM MM 2022 · 7 citations
- SAT3D: Image-driven Semantic Attribute Transfer in 3DZhijun Zhai, Zengmao Wang, Xiaoxiao Long, Kaixuan Zhou et al.ACM MM 2024
- Conceptual and Hierarchical Latent Space Decomposition for Face EditingSavas Özkan, Mete Özay, Tom RobinsonICCV 2023 · 3 citations
- Controllable 3D Face Generation with Conditional Style Code DiffusionXiaolong Shen, Jianxin Ma, Chang Zhou, Zongxin YangAAAI 2024 · 19 citations
- Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image ManipulationXiwen Wei, Zhen Xu, Cheng Liu, Si Wu et al.CVPR 2023
