ProsodyTalker: 3D Visual Speech Animation via Prosody Decomposition
Zonglin Li, Xiaoqian Lv, Qinglin Liu, Quanling Meng, Xin Sun, Shengping Zhang
摘要
Most existing 3D visual speech animation methods synthesize lip movements synchronized with speech, which however neglect head poses and therefore degrade the animation realism. The animation of head poses presents two primary challenges: (1) the intricate mapping between speech and head poses remains poorly understood and (2) the absence of 4D face datasets featuring realistic head poses. Inspired by prosody decomposition in speech processing, we discern that head movements correlate with the fundamental frequency (F0) of speech prosody, while lip movements align with the language content. These observations motivate us to propose a novel framework, dubbed ProsodyTalker, that concurrently synthesizes lip and head movements, grounded in the principles of prosody decomposition. The core idea is first to adopt information perturbation to explicitly decompose the speech prosody into pose-related F0 and lip-related language content. Then, an autoregressive content-oriented fusion decoder is employed to enhance lip synchronization in the synthesized facial sequences. To synthesize head poses, we design a transformer-based variational autoencoder to learn a latent distribution of facial sequences and propose an F0-conditioned latent diffusion model to establish a probabilistic mapping from F0 to pose-related latent codes. Furthermore, we contribute a large-scale 4D face dataset containing bunches of variations in identities, head poses and facial motions. Extensive experiments show that our method achieves more realistic animation than state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 被引用 534 次
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre 等ICCV 2021 · 被引用 272 次
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang 等CVPR 2022 · 被引用 218 次
相关 Paper
- Towards High-fidelity 3D Talking Avatar with Personalized Dynamic TextureXuanchen Li, Jianyu Wang, Yuhao Cheng, Yikun Zeng 等CVPR 2025
- CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorJinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun 等CVPR 2023
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja 等ACM MM 2025 · 被引用 3 次
- PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality AlignmentBin Wang, Yang Xu, Huan Zhao, Hao Zhang 等ACM MM 2025
- DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion AutoencoderChenpeng Du, Qi Chen, Tianyu He, Xu Tan 等ACM MM 2023 · 被引用 36 次
