Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
Qingcheng Zhao, Pengyu Long, Qixuan Zhang, Dafei Qin, Han Liang, Longwen Zhang, Yingliang Zhang, Jingyi Yu, Lan Xu
Abstract
The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism and a lack of flexible conditioning. We address this challenge through a trilogy. We first introduce Generalized Neural Parametric Facial Asset (GNPFA), an efficient variational auto-encoder mapping facial geometry and images to a highly generalized expression latent space, decoupling expressions and identities. Then, we utilize GNPFA to extract high-quality expressions and accurate head poses from a large array of videos. This presents the M2F-D dataset, a large, diverse, and scan-level co-speech 3D facial animation dataset with well-annotated emotion and style labels. Finally, we propose Media2Face, a diffusion model in GNPFA latent space for co-speech facial animation generation, accepting rich multi-modality guidances from audio, text, and image. Extensive experiments demonstrate that our model not only achieves high fidelity in facial animation synthesis but also broadens the scope of expressiveness and style adaptability in 3D facial animation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecae4e34-f1a1-45c0-9326-f67d3c555fb7Cited by top-tier papers17
- MMHead: Towards Fine-grained Multi-modal 3D Facial AnimationSijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan et al.ACM MM 2024 · 17 citations
- Examining the Effects of Immersive and Non-Immersive Presenter Modalities on Engagement and Social Interaction in Co-located Augmented PresentationsMatt Gottsacker, Mengyu Chen, David Saffo, Feiyu Lu et al.CHI 2025 · 6 citations
- InstructAvatar: Text-Guided Emotion and Motion Control for Avatar GenerationYuchi Wang, Junliang Guo, Jianhong Bai, Runyi Yu et al.AAAI 2025 · 5 citations
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme CasesShuai Tan, Bill Gong, Bin Ji, Ye PanICCV 2025 · 3 citations
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu et al.CVPR 2022 · 1,425 citations
Related papers
- DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion ModelsZhiyao Sun, Tian Lv, Sheng Ye, Matthieu Gaetan Lin et al.SIGGRAPH 2024 · 81 citations
- MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label GenerationXiangdong Li, Ye Lou, Ao Gao, Wei Zhang et al.AAAI 2026
- FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head ModelsShivangi Aneja, Justus Thies, Angela Dai, Matthias NießnerCVPR 2024
- StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion ModelYifan Yang, Zhi Cen, Sida Peng, Xiangwei Chen et al.AAAI 2026 · 1 citation
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
