TediGAN: Text-Guided Diverse Face Image Generation and Manipulation
Weihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan Wu
Abstract
In this work, we propose TediGAN, a novel framework for multi-modal image generation and manipulation with textual descriptions. The proposed method consists of three components: StyleGAN inversion module, visual-linguistic similarity learning, and instance-level optimization. The inversion module maps real images to the latent space of a well-trained StyleGAN. The visual-linguistic similarity learns the text-image matching by mapping the image and text into a common embedding space. The instancelevel optimization is for identity preservation in manipulation. Our model can produce diverse and high-quality images with an unprecedented resolution at 1024 2 . Using a control mechanism based on style-mixing, our Tedi-GAN inherently supports image synthesis with multi-modal inputs, such as sketches or semantic labels, with or without instance guidance. To facilitate text-guided multimodal synthesis, we propose the Multi-Modal CelebA-HQ, a large-scale dataset consisting of real face images and corresponding semantic segmentation map, sketch, and textual descriptions. Extensive experiments on the introduced dataset demonstrate the superior performance of our proposed method. Code and data are available at https://github.com/weihaox/TediGAN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db9dcaf3-8253-4f19-9302-15bd00162567Cited by top-tier papers118
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan et al.ICCV 2023 · 770 citations
- ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image GenerationYuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai et al.ICCV 2023 · 469 citations
- DiffusionCLIP: Text-Guided Diffusion Models for Robust Image ManipulationGwanghyun Kim, Taesung Kwon, Jong Chul YeCVPR 2022 · 458 citations
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman et al.ICLR 2023 · 361 citations
- CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance FieldsCan Wang, Menglei Chai, Mingming He, Dongdong Chen et al.CVPR 2022 · 313 citations
Builds on7
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- Seeing What a GAN Cannot GenerateDavid Bau, Jun-Yan Zhu, Jonas Wulff, William S. Peebles et al.ICCV 2019 · 342 citations
- Interactive Sketch & Fill: Multiclass Sketch-to-Image TranslationArnab Ghosh, Richard Zhang, Puneet K. Dokania, Oliver Wang et al.ICCV 2019 · 148 citations
- Lightweight Generative Adversarial Networks for Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Philip H. S. Torr, Thomas LukasiewiczNeurIPS 2020 · 76 citations
- Encoding in Style: A StyleGAN Encoder for Image-to-Image TranslationElad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan et al.CVPR 2021
Related papers
- Towards Open-Ended Text-to-Face Generation, Combination and ManipulationJun Peng, Han Pan, Yiyi Zhou, Jing He et al.ACM MM 2022 · 7 citations
- Controllable 3D Face Generation with Conditional Style Code DiffusionXiaolong Shen, Jianxin Ma, Chang Zhou, Zongxin YangAAAI 2024 · 19 citations
- Text-Conditional Attribute Alignment Across Latent Spaces for 3D Controllable Face Image SynthesisFeifan Xu, Rui Li, Si Wu, Yong Xu et al.CVPR 2024
- Diffusion-Driven GAN Inversion for Multi-Modal Face Image GenerationJihyun Kim, Changjae Oh, Hoseok Do, Soohyun Kim et al.CVPR 2024
- Self-Supervised Geometry-Aware Encoder for Style-Based 3D GAN InversionYushi Lan, Xuyi Meng, Shuai Yang, Chen Change Loy et al.CVPR 2023
