MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation
Huaize Liu, Wenzhang Sun, Donglin Di, Shibo Sun, Jiahui Yang, Changqing Zou, Hujun Bao
摘要
The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current methods face fundamental challenges: a) the absence of frameworks for modeling single basic emotional expressions, which restricts the generation of complex emotions such as compound emotions; b) the lack of comprehensive datasets rich in human emotional expressions, which limits the potential of models. To address these challenges, we propose the following innovations: 1) the Mixture of Emotion Experts (MoEE) model, which decouples six fundamental emotions to enable the precise synthesis of both singular and compound emotional states; 2) the DH-FaceEmoVid-150 dataset, specifically curated to include six prevalent human emotional expressions as well as four types of compound emotions, thereby expanding the training potential of emotion-driven models. Furthermore, to enhance the flexibility of emotion control, we propose an emotion-to-latents module that leverages multimodal inputs, aligning diverse control signals-such as audio, text, and labels-to ensure more varied control inputs as well as the ability to control emotions using audio alone. Through extensive quantitative and qualitative evaluations, we demonstrate that the MoEE framework, in conjunction with the DH-FaceEmoVid-150 dataset, excels in generating complex emotional expressions and nuanced facial details, setting a new benchmark in the field. These datasets will be publicly released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake DetectionYoungseo Kim, Kwan Yun, Seokhyeon Hong, Sihun Cha 等CVPR 2026 · 被引用 2 次
- Cross-Modal Emotion Transfer for Emotion Editing in Talking Face VideoChanhyuk Choi, Taesoo Kim, Donggyu Lee, Siyeol Jung 等CVPR 2026 · 被引用 1 次
- Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive ModelsLongtao Jiang, Jie Huang, Mingfei Han, Lei Chen 等AAAI 2026
- EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and GenerationZongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He 等AAAI 2026
它引用的顶会 Paper21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- From Sparse to Soft Mixtures of ExpertsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Neil HoulsbyICLR 2024 · 被引用 264 次
相关 Paper
- Expressive Talking AvatarsYe Pan, Shuai Tan, Shengran Cheng, Qunfen Lin 等IEEE VR 2024 · 被引用 20 次
- ECAvatar: 3D Avatar Facial Animation with Controllable Identity and EmotionMinjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun 等ACM MM 2024 · 被引用 4 次
- PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face GenerationBaiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu 等CVPR 2026 · 被引用 8 次
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja 等ACM MM 2025 · 被引用 3 次
- EmoFace: Audio-driven Emotional 3D Face AnimationChang Liu, Qunfen Lin, Zijiao Zeng, Ye PanIEEE VR 2024 · 被引用 16 次
