xADA: Controllable and Expressive Audio-Driven Animation
Sarah Taylor, Salvador Medina, Jonathan Windle, Erica Alcusa Sáez, Iain A. Matthews
Abstract
Fig. 1. 𝑥ADA generates high fidelity animation of the face, head, and tongue from speech audio, with accurate blinking and expression. 𝑥ADA is fully automatic, and can accurately animate a diverse range of speech and non-verbal sounds.
We introduce 𝑥ADA, a generative model for creating expressive and realistic animation of the face, tongue, and head directly from speech audio. Our approach leverages the pretrained Whisper audio encoder to extract rich speech features which are decoded into face and head animation using a series of gated recurrent unit (GRU) networks. The generated animation maps directly onto MetaHuman compatible rig controls enabling seamless integration into industry-standard content creation pipelines. 𝑥ADA operates fully automatically, with an option for users to override the detected emotion and/or blink timings. 𝑥ADA generalizes across languages, and voice styles, and can animate non-verbal sounds. Quantitative evaluation and a user study demonstrate that 𝑥ADA produces state-of-the-art animation with high realism, frequently indistinguishable from ground truth performance. Additionally, we outline a comprehensive data capture protocol designed to collect an extensive range of speech and non-verbal sounds for training animation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
Related papers
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- Speech Driven Tongue AnimationSalvador Medina, Denis Tomè, Carsten Stoll, Mark Tiede et al.CVPR 2022 · 14 citations
- Semi-supervised Speech-driven 3D Facial Animation via Cross-modal EncodingPeiji Yang, Huawei Wei, Yicheng Zhong, Zhisheng WangICCV 2023 · 1 citation
- VASA-Rig: Audio-Driven 3D Facial Animation with 'Live' Mood Dynamics in Virtual RealityYe Pan, Chang Liu, Sicheng Xu, Shuai Tan et al.IEEE VR 2025 · 5 citations
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding et al.AAAI 2021 · 88 citations
