SPACE: Speech-driven Portrait Animation with Controllable Expression
Siddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle, Ming-Yu Liu
摘要
Animating portraits using speech has received growing attention in recent years, with various creative and practical use cases. An ideal generated video should have good lip sync with the audio, natural facial expressions and head motions, and high frame quality. In this work, we present SPACE, which uses speech and a single image to generate high-resolution, and expressive videos with realistic head pose, without requiring a driving video. It uses a multi-stage approach, combining the controllability of facial landmarks with the high-quality synthesis power of a pretrained face generator. SPACE also allows for the control of emotions and their intensities. Our method outperforms prior methods in objective metrics for image quality and facial motions and is strongly preferred by users in pair-wise comparisons. Please visit the project page to view the videos and to see more results: https://research.nvidia.com/labs/dir/space/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion AutoencoderChenpeng Du, Qi Chen, Tianyu He, Xu Tan 等ACM MM 2023 · 被引用 36 次
- Expressive Talking AvatarsYe Pan, Shuai Tan, Shengran Cheng, Qunfen Lin 等IEEE VR 2024 · 被引用 20 次
- Speech-Driven 3D Face Animation with Composite and Regional Facial MovementsHaozhe Wu, Songtao Zhou, Jia Jia, Junliang Xing 等ACM MM 2023 · 被引用 20 次
- GaussianSpeech: Audio-Driven Personalized 3D Gaussian AvatarsShivangi Aneja, Artem Sevastopolsky, Tobias Kirschstein, Justus Thies 等ICCV 2025 · 被引用 8 次
- FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion ModelZiyu Yao, Xuxin Cheng, Zhiqi HuangACM MM 2024 · 被引用 5 次
它引用的顶会 Paper12
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang 等CVPR 2022 · 被引用 218 次
- MarioNETte: Few-Shot Face Reenactment Preserving Identity of Unseen TargetsSungjoo Ha, Martin Kersner, Beomsu Kim, Seokjun Seo 等AAAI 2020 · 被引用 184 次
- EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelXinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu 等SIGGRAPH 2022 · 被引用 150 次
- Efficient Facial Feature Learning with Wide Ensemble-Based Convolutional Neural NetworksHenrique Siqueira, Sven Magg, Stefan WermterAAAI 2020 · 被引用 136 次
相关 Paper
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang 等NeurIPS 2024 · 被引用 253 次
- That's What I Said: Fully-Controllable Talking Face GenerationYoungjoon Jang, Kyeongha Rho, Jong-Bin Woo, Hyeongkeun Lee 等ACM MM 2023 · 被引用 7 次
- Talking Head from Speech Audio using a Pre-trained Image GeneratorMohammed M. Alghamdi, He Wang, Andrew J. Bulpitt, David C. HoggACM MM 2022 · 被引用 25 次
- FLOAT: Generative Motion Latent Flow Matching for Audio-Driven Talking PortraitTaekyung Ki, Dongchan Min, Gyeongsu ChaeICCV 2025 · 被引用 6 次
- FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head ModelsShivangi Aneja, Justus Thies, Angela Dai, Matthias NießnerCVPR 2024
