Multi-caption Text-to-Face Synthesis: Dataset and Algorithm
Jianxin Sun, Qi Li, Weining Wang, Jian Zhao, Zhenan Sun
摘要
Text-to-Face synthesis with multiple captions is still an important yet less addressed problem because of the lack of effective algorithms and large-scale datasets. We accordingly propose a Semantic Embedding and Attention (SEA-T2F) network that allows multiple captions as input to generate highly semantically related face images. With a novel Sentence Features Injection Module, SEA-T2F can integrate any number of captions into the network. In addition, an attention mechanism named Attention for Multiple Captions is proposed to fuse multiple word features and synthesize fine-grained details. Considering text-to-face generation is an ill-posed problem, we also introduce an attribute loss to guide the network to generate sentence-related attributes. Existing datasets for text-to-face are either too small or roughly generated according to attribute labels, which is not enough to train deep learning based methods to synthesize natural face images. Therefore, we build a large-scale dataset named CelebAText-HQ, in which each image is manually annotated with 10 captions. Extensive experiments demonstrate the effectiveness of our algorithm.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- Improving Subgraph Recognition with Variational Graph Information BottleneckJunchi Yu, Jie Cao, Ran HeCVPR 2022 · 被引用 56 次
- AnyFace: Free-style Text-to-Face Synthesis and ManipulationJianxin Sun, Qiyao Deng, Qi Li, Muyi Sun 等CVPR 2022 · 被引用 46 次
- Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only ImagesCuican Yu, Guansong Lu, Yihan Zeng, Jian Sun 等ICCV 2023 · 被引用 20 次
- Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationXintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding 等ACM MM 2022 · 被引用 17 次
- Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric RegularizationJinlu Zhang, Yiyi Zhou, Qiancheng Zheng, Xiaoxiong Du 等ICML 2024 · 被引用 9 次
相关 Paper
- RiFeGAN: Rich Feature Generation for Text-to-Image Synthesis From Prior KnowledgeJun Cheng, Fuxiang Wu, Yanling Tian, Lei Wang 等CVPR 2020
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 被引用 13 次
- TediGAN: Text-Guided Diverse Face Image Generation and ManipulationWeihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan WuCVPR 2021
- SDGAN: Disentangling Semantic Manipulation for Facial Attribute EditingWenmin Huang, Weiqi Luo, Jiwu Huang, Xiaochun CaoAAAI 2024 · 被引用 20 次
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding 等AAAI 2021 · 被引用 88 次
