DEITalk: Speech-Driven 3D Facial Animation with Dynamic Emotional Intensity Modeling
Kang Shen, Haifeng Xia, Guangxing Geng, Guangyue Geng, Siyu Xia, Zhengming Ding
Abstract
Speech-driven 3D facial animation aims to synthesize 3D talking head animations with precise lip movements and rich stylistic expressions. However, existing methods exhibit two limitations: 1) they mostly focused on emotionless facial animation modeling, neglecting the importance of emotional expression, due to the lack of high-quality 3D emotional talking head datasets, and 2) several latest works treated emotional intensity as a global controllable parameter, akin to emotional or speaker style, leading to over-smoothed emotional expressions in their outcomes. To address these challenges, we first collect a 3D talking head dataset comprising five emotional styles with a set of coefficients based on the MetaHuman character model and then propose an end-to-end deep neural network, DEITalk, which conditions on speech and emotional style labels to generate realistic facial animation with dynamic expressions. To model emotional saliency variations in long-term audio contexts, we design a dynamic emotional intensity (DEI) modeling module and a dynamic positional encoding (DPE) strategy. The former extracts implicit representations of emotional intensity from speech features and utilizes them as local (high temporal frequency) emotional supervision, whereas the latter offers abilities to generalize to longer speech sequences. Moreover, we introduce an emotion-guided feature fusion decoder and a four-way loss function to generate emotion-enhanced 3D facial animation with controllable emotional styles. Extensive experimental results demonstrate that our method outperforms existing state-of-the-art methods. Our video demo and dataset are available at https://github.com/KangShen-seu/DEITalk.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9b6ff63c-58f4-4e42-9acf-de14b2eac061Cited by top-tier papers2
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional StylesTianshun Han, Benjia Zhou, Ajian Liu, Yanyan Liang et al.ACM MM 2025 · 1 citation
Related papers
- DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face AnimationJisoo Kim, Jungbin Cho, Joonho Park, Soonmin Hwang et al.AAAI 2025 · 13 citations
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
- PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face GenerationBaiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu et al.CVPR 2026 · 8 citations
- EmoFace: Audio-driven Emotional 3D Face AnimationChang Liu, Qunfen Lin, Zijiao Zeng, Ye PanIEEE VR 2024 · 16 citations
- VASA-Rig: Audio-Driven 3D Facial Animation with 'Live' Mood Dynamics in Virtual RealityYe Pan, Chang Liu, Sicheng Xu, Shuai Tan et al.IEEE VR 2025 · 5 citations
