Lune

SIGGRAPH2025Top-tier venue

xADA: Controllable and Expressive Audio-Driven Animation

Sarah Taylor, Salvador Medina, Jonathan Windle, Erica Alcusa Sáez, Iain A. Matthews

2025Year
2Citations

Abstract

Fig. 1. 𝑥ADA generates high fidelity animation of the face, head, and tongue from speech audio, with accurate blinking and expression. 𝑥ADA is fully automatic, and can accurately animate a diverse range of speech and non-verbal sounds.

We introduce 𝑥ADA, a generative model for creating expressive and realistic animation of the face, tongue, and head directly from speech audio. Our approach leverages the pretrained Whisper audio encoder to extract rich speech features which are decoded into face and head animation using a series of gated recurrent unit (GRU) networks. The generated animation maps directly onto MetaHuman compatible rig controls enabling seamless integration into industry-standard content creation pipelines. 𝑥ADA operates fully automatically, with an option for users to override the detected emotion and/or blink timings. 𝑥ADA generalizes across languages, and voice styles, and can animate non-verbal sounds. Quantitative evaluation and a user study demonstrate that 𝑥ADA produces state-of-the-art animation with high realism, frequently indistinguishable from ground truth performance. Additionally, we outline a comprehensive data capture protocol designed to collect an extensive range of speech and non-verbal sounds for training animation models.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines