Talking Head Generation with Probabilistic Audio-to-Visual Diffusion Priors
Zhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang, Finn Wong, Baoyuan Wang
Abstract
We introduce a novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead sample all holistic lip-irrelevant facial motions (i.e. pose, expression, blink, gaze, etc.) to semantically match the input audio while still maintaining both the photo-realism of audio-lip synchronization and overall naturalness. This is achieved by our newly proposed audio-to-visual diffusion prior trained on top of the mapping between audio and non-lip representations. Thanks to the probabilistic nature of the diffusion prior, one big advantage of our framework is it can synthesize diverse facial motion sequences given the same audio clip, which is quite user-friendly for many real applications. Through comprehensive evaluations of public benchmarks, we conclude that (1) our diffusion prior outperforms auto-regressive prior significantly on all the concerned metrics; (2) our overall system is competitive with prior works in terms of audio-lip synchronization but can effectively sample rich and natural-looking lip-irrelevant facial motions while still semantically harmonized with the audio input.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisJiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou et al.ICCV 2023 · 127 citations
- HumanTOMATO: Text-aligned Whole-body Motion GenerationShunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin et al.ICML 2024 · 124 citations
- HumanMAC: Masked Motion Completion for Human Motion PredictionLing-Hao Chen, Jiawei Zhang, Yewen Li, Yiren Pang et al.ICCV 2023 · 106 citations
- GAIA: Zero-shot Talking Avatar GenerationTianyu He, Junliang Guo, Runyi Yu, Yuchi Wang et al.ICLR 2024 · 51 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionsZiqi Zhou, Weize Quan, Hailin Shi, Wei Li et al.AAAI 2025 · 1 citation
- DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking head Video GenerationHanbo Cheng, Limin Lin, Chenyu Liu, Pengcheng Xia et al.ICLR 2025
- FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute LearningChenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng et al.ICCV 2021 · 149 citations
- Towards Realistic Visual Dubbing with Heterogeneous SourcesTianyi Xie, Liucheng Liao, Cheng Bi, Benlai Tang et al.ACM MM 2021 · 34 citations
- Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head SynthesisDuomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum et al.CVPR 2023
