EcoFace: Audio-Visual Emotional Co-Disentanglement Speech-Driven 3D Talking Face Generation
Jiajian Xie, Shengyu Zhang, Mengze Li, Chengfei Lv, Zhou Zhao, Fei Wu
Abstract
Speech-driven 3D facial animation has attracted significant attention due to its wide range of applications in animation production and virtual reality. Recent research has explored speech-emotion disentanglement to enhance facial expressions rather than manually assigning emotions. However, this approach face issues such as feature confusion, emotions weakening and mean-face. To address these issues, we present EcoFace, a framework that (1) proposes a novel collaboration objective to provide an explicit signal for emotion representation learning from the speaker's expressive movements and produced sounds, constructing an audiovisual joint and coordinated emotion space that is independent of speech content.
(2) constructs a universal facial motion distribution space determined by speech features and implement speaker-aware generation. Extensive experiments show that our method achieves more generalized and emotionally realistic talking face generation compared to previous methods 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on15
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 460 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
Related papers
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
- ECAvatar: 3D Avatar Facial Animation with Controllable Identity and EmotionMinjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun et al.ACM MM 2024 · 4 citations
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial AnimationHui Fu, Zeqing Wang, Ke Gong, Keze Wang et al.AAAI 2024 · 23 citations
- DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face AnimationJisoo Kim, Jungbin Cho, Joonho Park, Soonmin Hwang et al.AAAI 2025 · 13 citations
