DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation
Jisoo Kim, Jungbin Cho, Joonho Park, Soonmin Hwang, Da Eun Kim, Geon Kim, Youngjae Yu
Abstract
Speech-driven 3D facial animation has garnered lots of attention thanks to its broad range of applications. Despite recent advancements in achieving realistic lip motion, current methods fail to capture the nuanced emotional undertones conveyed through speech and produce monotonous facial motion. These limitations result in blunt and repetitive facial animations, reducing user engagement and hindering their applicability. To address these challenges, we introduce DEEPTalk, a novel approach that generates diverse and emotionally rich 3D facial expressions directly from speech inputs. To achieve this, we first train DEE (Dynamic Emotion Embedding), which employs probabilistic contrastive learning to forge a joint emotion embedding space for both speech and facial motion. This probabilistic framework captures the uncertainty in interpreting emotions from speech and facial motion, enabling the derivation of emotion vectors from its multifaceted space. Moreover, to generate dynamic facial motion, we design TH-VQVAE (Temporally Hierarchical VQ-VAE) as an expressive and robust motion prior overcoming limitations of VAEs and VQ-VAEs. Utilizing these strong priors, we develop DEEPTalk, a talking head generator that non-autoregressively predicts codebook indices to create dynamic facial motion, incorporating a novel emotion consistency loss. Extensive experiments on various datasets demonstrate the effectiveness of our approach in creating diverse, emotionally expressive talking faces that maintain accurate lip-sync. Our project page is available at https://whwjdqls.github.io/deeptalk.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8a43b14-f462-481c-bb3e-b02b77ce58abCited by top-tier papers3
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional StylesTianshun Han, Benjia Zhou, Ajian Liu, Yanyan Liang et al.ACM MM 2025 · 1 citation
- EmoDiffTalk: Emotion-aware Diffusion for Editable 3D Gaussian Talking HeadChang Liu, Tianjiao Jing, Chengcheng Ma, Xuanqi Zhou et al.CVPR 2026 · 1 citation
Builds on19
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang et al.CVPR 2022 · 218 citations
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
- EMOCA: Emotion Driven Monocular Face Capture and AnimationRadek Danecek, Michael J. Black, Timo BolkartCVPR 2022 · 180 citations
- EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelXinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu et al.SIGGRAPH 2022 · 150 citations
Related papers
- DEITalk: Speech-Driven 3D Facial Animation with Dynamic Emotional Intensity ModelingKang Shen, Haifeng Xia, Guangxing Geng, Guangyue Geng et al.ACM MM 2024 · 6 citations
- Towards Variable and Coordinated Holistic Co-Speech Motion GenerationYifei Liu, Qiong Cao, Yandong Wen, Huaiguang Jiang et al.CVPR 2024
- CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorJinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun et al.CVPR 2023
- FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and QuantizationShuai Tan, Bin Ji, Ye PanCVPR 2024 · 17 citations
- Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait GenerationWeipeng Tan, Chuming Lin, Chengming Xu, FeiFan Xu et al.ACM MM 2025 · 4 citations
