High-Fidelity Generalized Emotional Talking Face Generation with Multi-Modal Emotion Space Learning
Chao Xu, Junwei Zhu, Jiangning Zhang, Yue Han, Wenqing Chu, Ying Tai, Chengjie Wang, Zhifeng Xie, Yong Liu
摘要
Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing to handle unseen emotion styles due to limited semantics. They either ignore the one-shot setting or the quality of generated faces. In this paper, we propose a more flexible and generalized framework. Specifically, we supplement the emotion style in text prompts and use an Aligned Multi-modal Emotion encoder to embed the text, image, and audio emotion modality into a unified space, which inherits rich semantic prior from CLIP. Consequently, effective multi-modal emotion space learning helps our method support arbitrary emotion modality during testing and could generalize to unseen emotion styles. Besides, an Emotionaware Audio-to-3DMM Convertor is proposed to connect the emotion condition and the audio sequence to structural representation. A followed style-based High-fidelity Emotional Face generator is designed to generate arbitrary high-resolution realistic identities. Our texture generator hierarchically learns flow fields and animated faces in a residual manner. Extensive experiments demonstrate the flexibility and generalization of our method in emotion control and the effectiveness of high-quality face synthesis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- PortraitBooth: A Versatile Portrait Model for Fast Identity-Preserved PersonalizationXu Peng, Junwei Zhu, Boyuan Jiang, Ying Tai 等CVPR 2024 · 被引用 28 次
- ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion ModelingQuanwei Yang, Jiazhi Guan, Kaisiyuan Wang, Lingyun Yu 等NeurIPS 2024 · 被引用 21 次
- FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-Pose, and Facial Expression FeaturesAndre Rochow, Max Schwarz, Sven BehnkeCVPR 2024 · 被引用 17 次
- FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and QuantizationShuai Tan, Bin Ji, Ye PanCVPR 2024 · 被引用 17 次
- MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion TransformerShurong Yang, Huadong Li, Juhao Wu, Minhao Jing 等AAAI 2025 · 被引用 12 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
相关 Paper
- MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech SynthesisWenhao Guan, Yishuang Li, Tao Li, Hukai Huang 等AAAI 2024 · 被引用 25 次
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja 等ACM MM 2025 · 被引用 3 次
- Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art StyleShuai Tan, Bin Ji, Ye PanAAAI 2024
- PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face GenerationBaiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu 等CVPR 2026 · 被引用 8 次
- EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelXinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu 等SIGGRAPH 2022 · 被引用 150 次
