Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art Style
Shuai Tan, Bin Ji, Ye Pan
Abstract
Although automatically animating audio-driven talking heads has recently received growing interest, previous efforts have mainly concentrated on achieving lip synchronization with the audio, neglecting two crucial elements for generating expressive videos: emotion style and art style. In this paper, we present an innovative audio-driven talking face generation method called Style 2 Talker. It involves two stylized stages, namely Style-E and Style-A, which integrate text-controlled emotion style and picture-controlled art style into the final output. In order to prepare the scarce emotional text descriptions corresponding to the videos, we propose a labor-free paradigm that employs large-scale pretrained models to automatically annotate emotional text labels for existing audiovisual datasets. Incorporating the synthetic emotion texts, the Style-E stage utilizes a large-scale CLIP model to extract emotion representations, which are combined with the audio, serving as the condition for an efficient latent diffusion model designed to produce emotional motion coefficients of a 3DMM model. Moving on to the Style-A stage, we develop a coefficient-driven motion generator and an art-specific style path embedded in the well-known StyleGAN. This allows us to synthesize high-resolution artistically stylized talking head videos using the generated emotional motion coefficients and an art style source picture. Moreover, to better preserve image details and avoid artifacts, we provide StyleGAN with the multi-scale content features extracted from the identity image and refine its intermediate feature maps by the designed content encoder and refinement network, respectively. Extensive experimental results demonstrate our method outperforms existing state-of-the-art methods in terms of audio-lip synchronization and performance of both emotion style and art style.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and QuantizationShuai Tan, Bin Ji, Ye PanCVPR 2024 · 17 citations
- An Efficient and Harmonized Framework for Balanced Cross-Domain Feature IntegrationShaoxu Li, Ye PanAAAI 2026 · 15 citations
- OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style MimickingZhongjian Wang, Peng Zhang, Jinwei Qi, Yuan Wang et al.NeurIPS 2025 · 12 citations
- SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local EditingLingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu et al.ACM MM 2024 · 10 citations
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video GenerationShuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang et al.CVPR 2026 · 9 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
Related papers
- High-Fidelity Generalized Emotional Talking Face Generation with Multi-Modal Emotion Space LearningChao Xu, Junwei Zhu, Jiangning Zhang, Yue Han et al.CVPR 2023
- Model See Model Do: Speech-Driven Facial Animation with Style ControlYifang Pan, Karan Singh, Luiz Gustavo HafemannSIGGRAPH 2025 · 2 citations
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- Talking Head from Speech Audio using a Pre-trained Image GeneratorMohammed M. Alghamdi, He Wang, Andrew J. Bulpitt, David C. HoggACM MM 2022 · 25 citations
- PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face GenerationBaiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu et al.CVPR 2026 · 8 citations
