DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion Models
Zhiyao Sun, Tian Lv, Sheng Ye, Matthieu Gaetan Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, Yong-Jin Liu
摘要
The generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion. However, existing methods either employ a deterministic model for speech-to-motion mapping or encode the style using a one-hot encoding scheme. Notably, the one-hot encoding approach fails to capture the complexity of the style and thus limits generalization ability. In this paper, we propose DiffPoseTalk, a generative framework based on the diffusion model combined with a style encoder that extracts style embeddings from short reference videos. During inference, we employ classifier-free guidance to guide the generation process based on the speech and style. In particular, our style includes the generation of head poses, thereby enhancing user perception. Additionally, we address the shortage of scanned 3D talking face data by training our model on reconstructed 3DMM parameters from a high-quality, in-the-wild audio-visual dataset. Extensive experiments and user study demonstrate that our approach outperforms state-of-the-art methods. The code and dataset are at https://diffposetalk.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang 等NeurIPS 2024 · 被引用 253 次
- Media2Face: Co-speech Facial Animation Generation With Multi-Modality GuidanceQingcheng Zhao, Pengyu Long, Qixuan Zhang, Dafei Qin 等SIGGRAPH 2024 · 被引用 40 次
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human AvatarsZhiyao Sun, Ziqiao Peng, Yifeng Ma, Yi Chen 等CVPR 2026 · 被引用 26 次
- MMHead: Towards Fine-grained Multi-modal 3D Facial AnimationSijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan 等ACM MM 2024 · 被引用 17 次
- UniLS: End-to-End Audio-Driven Avatars for Unified Listening and SpeakingXuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu 等CVPR 2026 · 被引用 12 次
它引用的顶会 Paper16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
相关 Paper
- Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face SynthesisHaozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou 等ACM MM 2021 · 被引用 50 次
- Model See Model Do: Speech-Driven Facial Animation with Style ControlYifang Pan, Karan Singh, Luiz Gustavo HafemannSIGGRAPH 2025 · 被引用 2 次
- DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion AutoencoderChenpeng Du, Qi Chen, Tianyu He, Xu Tan 等ACM MM 2023 · 被引用 36 次
- CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorJinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun 等CVPR 2023
- DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face AnimationJisoo Kim, Jungbin Cho, Joonho Park, Soonmin Hwang 等AAAI 2025 · 被引用 13 次
