VASA-Rig: Audio-Driven 3D Facial Animation with 'Live' Mood Dynamics in Virtual Reality
Ye Pan, Chang Liu, Sicheng Xu, Shuai Tan, Jiaolong Yang
Abstract
Audio-driven 3D facial animation is crucial for enhancing the metaverse's realism, immersion, and interactivity. While most existing methods focus on generating highly realistic and lively 2D talking head videos by leveraging extensive 2D video datasets these approaches work in pixel space and are not easily adaptable to 3D environments. We present VASA-Rig, which has achieved a significant advancement in the realism of lip-audio synchronization, facial dynamics, and head movements. In particular, we introduce a novel rig parameter-based emotional talking face dataset and propose the Latents2Rig model, which facilitates the transformation of 2D facial animations into 3D. Unlike mesh-based models, VASA-Rig outputs rig parameters, instantiated in this paper as 174 Metahuman rig parameters, making it more suitable for integration into industry-standard pipelines. Extensive experimental results demonstrate that our approach significantly outperforms existing state-of-the-art methods in terms of both realism and accuracy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 42069a26-135d-41f3-8b6f-ba0d71bb7feeCited by top-tier papers3
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme CasesShuai Tan, Bill Gong, Bin Ji, Ye PanICCV 2025 · 3 citations
- HiFi-Mesh: High-Fidelity Efficient 3D Mesh Generation via Compact Autoregressive DependenceYanfeng Li, Tao Tan, Qinquan Gao, Zhiwen Cao et al.AAAI 2026
Related papers
- VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single ImageSicheng Xu, Guojun Chen, Jiaolong Yang, Yizhong Zhang et al.NeurIPS 2025 · 5 citations
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- EmoFace: Audio-driven Emotional 3D Face AnimationChang Liu, Qunfen Lin, Zijiao Zeng, Ye PanIEEE VR 2024 · 16 citations
- DEITalk: Speech-Driven 3D Facial Animation with Dynamic Emotional Intensity ModelingKang Shen, Haifeng Xia, Guangxing Geng, Guangyue Geng et al.ACM MM 2024 · 6 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
