VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
Sicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, Baining Guo
摘要
We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip movements that are exquisitely synchronized with the audio, but also producing a large spectrum of facial nuances and natural head motions that contribute to the perception of authenticity and liveliness. The core innovations include a holistic facial dynamics and head movement generation model that works in a face latent space, and the development of such an expressive and disentangled face latent space using videos. Through extensive experiments including evaluation on a set of new metrics, we show that our method significantly outperforms previous methods along various dimensions comprehensively. Our method not only delivers high video quality with realistic facial and head dynamics but also supports the online generation of 512x512 videos at up to 40 FPS with negligible starting latency. It paves the way for real-time engagements with lifelike avatars that emulate human conversational behaviors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper70
- SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human GenerationYouliang Zhang, Zhaoyang Li, Duomin Wang, jiahe zhang 等ICLR 2026 · 被引用 30 次
- Instilling an Active Mind in Avatars via Cognitive SimulationJianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang 等ICLR 2026 · 被引用 26 次
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human AvatarsZhiyao Sun, Ziqiao Peng, Yifeng Ma, Yi Chen 等CVPR 2026 · 被引用 26 次
- EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human AnimationRang Meng, Yan Wang, Weipeng Wu, Ruobing Zheng 等AAAI 2026 · 被引用 24 次
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio ConditionsZhenzhi Wang, Jiaqi Yang, Jianwen Jiang, Chao Liang 等ICLR 2026 · 被引用 19 次
它引用的顶会 Paper38
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
相关 Paper
- VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single ImageSicheng Xu, Guojun Chen, Jiaolong Yang, Yizhong Zhang 等NeurIPS 2025 · 被引用 5 次
- VASA-Rig: Audio-Driven 3D Facial Animation with 'Live' Mood Dynamics in Virtual RealityYe Pan, Chang Liu, Sicheng Xu, Shuai Tan 等IEEE VR 2025 · 被引用 5 次
- That's What I Said: Fully-Controllable Talking Face GenerationYoungjoon Jang, Kyeongha Rho, Jong-Bin Woo, Hyeongkeun Lee 等ACM MM 2023 · 被引用 7 次
- SPACE: Speech-driven Portrait Animation with Controllable ExpressionSiddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle 等ICCV 2023 · 被引用 58 次
- Synthesizing Photorealistic Virtual Humans Through Cross-Modal DisentanglementSiddarth Ravichandran, Ondrej Texler, Dimitar Dinev, Hyun Jae KangCVPR 2023
