Emotional Speech-Driven 3D Body Animation via Disentangled Latent Diffusion
Kiran Chhatre, Radek Danecek, Nikos Athanasiou, Giorgio Becherini, Christopher Peters, Michael J. Black, Timo Bolkart
摘要
Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output animations from speech without control over the expressed emotion. To address this limitation, we present AMUSE, an emotional speech-driven body animation model based on latent diffusion. Our observation is that content (i.e., gestures related to speech rhythm and word utterances), emotion, and personal style are separable. To account for this, AMUSE maps the driving audio to three disentangled latent vectors: one for content, one for emotion, and one for personal style. A latent diffusion model, trained to generate gesture motion sequences, is then conditioned on these latent vectors. Once trained, AMUSE synthesizes 3D human gestures directly from speech with control over the expressed emotions and style by combining the content from the driving speech with the emotion and style of another speech sequence. Randomly sampling the noise of the diffusion model further generates variations of the gesture with the same emotional expressivity. Qualitative, quantitative, and perceptual evaluations demonstrate that AMUSE outputs realistic gesture sequences. Compared to the state of the art, the generated gestures are better synchronized with the speech content, and better represent the emotion expressed by the input speech. Our code is available at amuse.is.tue.mpg.de.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space ModelsZunnan Xu, Yukang Lin, Haonan Han, Sicheng Yang 等NeurIPS 2024 · 被引用 62 次
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture ModelingHaiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng 等CVPR 2024 · 被引用 55 次
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao 等SIGGRAPH 2024 · 被引用 39 次
- Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion GenerationBohong Chen, Yumeng Li, Yao-Xiang Ding, Tianjia Shao 等ACM MM 2024 · 被引用 26 次
- SemTalk: Holistic Co-Speech Motion Generation with Frame-Level Semantic EmphasisXiangyue Zhang, Jianfang Li, Jiaxu Zhang, Ziqiang Dang 等ICCV 2025 · 被引用 12 次
它引用的顶会 Paper42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- Model See Model Do: Speech-Driven Facial Animation with Style ControlYifang Pan, Karan Singh, Luiz Gustavo HafemannSIGGRAPH 2025 · 被引用 2 次
- DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from SpeechYongkang Cheng, Shaoli Huang, Xuelin Chen, Jifeng Ning 等AAAI 2025 · 被引用 3 次
- Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional ControlZunnan Xu, Yachao Zhang, Sicheng Yang, Ronghui Li 等AAAI 2024 · 被引用 20 次
- ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Ikhsanul Habibie, Lucia Donatelli 等CVPR 2024 · 被引用 15 次
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion ModelsSimon Alexanderson, Rajmund Nagy, Jonas Beskow, Gustav Eje HenterSIGGRAPH 2023 · 被引用 191 次
