Multi-caption Text-to-Face Synthesis: Dataset and Algorithm
Jianxin Sun, Qi Li, Weining Wang, Jian Zhao, Zhenan Sun
Abstract
Text-to-Face synthesis with multiple captions is still an important yet less addressed problem because of the lack of effective algorithms and large-scale datasets. We accordingly propose a Semantic Embedding and Attention (SEA-T2F) network that allows multiple captions as input to generate highly semantically related face images. With a novel Sentence Features Injection Module, SEA-T2F can integrate any number of captions into the network. In addition, an attention mechanism named Attention for Multiple Captions is proposed to fuse multiple word features and synthesize fine-grained details. Considering text-to-face generation is an ill-posed problem, we also introduce an attribute loss to guide the network to generate sentence-related attributes. Existing datasets for text-to-face are either too small or roughly generated according to attribute labels, which is not enough to train deep learning based methods to synthesize natural face images. Therefore, we build a large-scale dataset named CelebAText-HQ, in which each image is manually annotated with 10 captions. Extensive experiments demonstrate the effectiveness of our algorithm.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers7
- Improving Subgraph Recognition with Variational Graph Information BottleneckJunchi Yu, Jie Cao, Ran HeCVPR 2022 · 56 citations
- AnyFace: Free-style Text-to-Face Synthesis and ManipulationJianxin Sun, Qiyao Deng, Qi Li, Muyi Sun et al.CVPR 2022 · 46 citations
- Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only ImagesCuican Yu, Guansong Lu, Yihan Zeng, Jian Sun et al.ICCV 2023 · 20 citations
- Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationXintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding et al.ACM MM 2022 · 17 citations
- Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric RegularizationJinlu Zhang, Yiyi Zhou, Qiancheng Zheng, Xiaoxiong Du et al.ICML 2024 · 9 citations
Related papers
- RiFeGAN: Rich Feature Generation for Text-to-Image Synthesis From Prior KnowledgeJun Cheng, Fuxiang Wu, Yanling Tian, Lei Wang et al.CVPR 2020
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 13 citations
- TediGAN: Text-Guided Diverse Face Image Generation and ManipulationWeihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan WuCVPR 2021
- SDGAN: Disentangling Semantic Manipulation for Facial Attribute EditingWenmin Huang, Weiqi Luo, Jiwu Huang, Xiaochun CaoAAAI 2024 · 20 citations
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding et al.AAAI 2021 · 88 citations
