AnyFace: Free-style Text-to-Face Synthesis and Manipulation
Jianxin Sun, Qiyao Deng, Qi Li, Muyi Sun, Min Ren, Zhenan Sun
摘要
Existing text-to-image synthesis methods generally are only applicable to words in the training dataset. However, human faces are so variable to be described with limited words. So this paper proposes the first free-style text-to-face method namely AnyFace enabling much wider open world applications such as metaverse, social media, cosmetics, forensics, etc. AnyFace has a novel two-stream framework for face image synthesis and manipulation given arbitrary descriptions of the human face. Specifically, one stream performs text-to-face generation and the other conducts face image reconstruction. Facial text and image features are extracted using the CLIP (Contrastive Language-Image Pre-training) encoders. And a collaborative Cross Modal Distillation (CMD) module is designed to align the linguistic and visual features across these two streams. Furthermore, a Diverse Triplet Loss (DT loss) is developed to model fine-grained features and improve facial diversity. Extensive experiments on Multi-modal CelebA-HQ and CelebAText-HQ demonstrate significant advantages of AnyFace over state-of-the-art methods. AnyFace can achieve high-quality, high-resolution, and high-diversity face synthesis and manipulation results without any constraints on the number and content of input captions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Text2Performer: Text-Driven Human Video GenerationYuming Jiang, Shuai Yang, Tong Liang Koh, Wayne Wu 等ICCV 2023 · 被引用 74 次
- CLIPVG: Text-Guided Image Manipulation Using Differentiable Vector GraphicsYiren Song, Xuning Shao, Kang Chen, Weidong Zhang 等AAAI 2023 · 被引用 50 次
- Pluralistic Aging Diffusion AutoencoderPeipei Li, Rui Wang, Huaibo Huang, Ran He 等ICCV 2023 · 被引用 25 次
- Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only ImagesCuican Yu, Guansong Lu, Yihan Zeng, Jian Sun 等ICCV 2023 · 被引用 20 次
- FaceComposer: A Unified Model for Versatile Facial Content CreationJiayu Wang, Kang Zhao, Yifeng Ma, Shiwei Zhang 等NeurIPS 2023 · 被引用 14 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
- Cycle-Consistent Inverse GAN for Text-to-Image SynthesisHao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan MiaoACM MM 2021 · 被引用 47 次
相关 Paper
- HairCLIP: Design Your Hair by Text and Reference ImageTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao 等CVPR 2022 · 被引用 94 次
- Towards Counterfactual Image Manipulation via CLIPYingchen Yu, Fangneng Zhan, Rongliang Wu, Jiahui Zhang 等ACM MM 2022 · 被引用 33 次
- Draw Your Art Dream: Diverse Digital Art Synthesis with Multimodal Guided DiffusionNisha Huang, Fan Tang, Weiming Dong, Changsheng XuACM MM 2022 · 被引用 49 次
- TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language NegativesMaitreya Patel, Abhiram Kusumba, Sheng Cheng, Changhoon Kim 等NeurIPS 2024 · 被引用 73 次
- ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion DesignXujie Zhang, Yu Sha, Michael C. Kampffmeyer, Zhenyu Xie 等ACM MM 2022 · 被引用 27 次
