xADA: Controllable and Expressive Audio-Driven Animation
Sarah Taylor, Salvador Medina, Jonathan Windle, Erica Alcusa Sáez, Iain A. Matthews
摘要
Fig. 1. 𝑥ADA generates high fidelity animation of the face, head, and tongue from speech audio, with accurate blinking and expression. 𝑥ADA is fully automatic, and can accurately animate a diverse range of speech and non-verbal sounds.
We introduce 𝑥ADA, a generative model for creating expressive and realistic animation of the face, tongue, and head directly from speech audio. Our approach leverages the pretrained Whisper audio encoder to extract rich speech features which are decoded into face and head animation using a series of gated recurrent unit (GRU) networks. The generated animation maps directly onto MetaHuman compatible rig controls enabling seamless integration into industry-standard content creation pipelines. 𝑥ADA operates fully automatically, with an option for users to override the detected emotion and/or blink timings. 𝑥ADA generalizes across languages, and voice styles, and can animate non-verbal sounds. Quantitative evaluation and a user study demonstrate that 𝑥ADA produces state-of-the-art animation with high realism, frequently indistinguishable from ground truth performance. Additionally, we outline a comprehensive data capture protocol designed to collect an extensive range of speech and non-verbal sounds for training animation models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre 等ICCV 2021 · 被引用 272 次
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang 等NeurIPS 2024 · 被引用 253 次
相关 Paper
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja 等ACM MM 2025 · 被引用 3 次
- Speech Driven Tongue AnimationSalvador Medina, Denis Tomè, Carsten Stoll, Mark Tiede 等CVPR 2022 · 被引用 14 次
- Semi-supervised Speech-driven 3D Facial Animation via Cross-modal EncodingPeiji Yang, Huawei Wei, Yicheng Zhong, Zhisheng WangICCV 2023 · 被引用 1 次
- VASA-Rig: Audio-Driven 3D Facial Animation with 'Live' Mood Dynamics in Virtual RealityYe Pan, Chang Liu, Sicheng Xu, Shuai Tan 等IEEE VR 2025 · 被引用 5 次
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding 等AAAI 2021 · 被引用 88 次
