A Motion Matching-based Framework for Controllable Gesture Synthesis from Speech
Ikhsanul Habibie, Mohamed A. Elgharib, Kripasindhu Sarkar, Ahsan Abdullah, Simbarashe Nyatsanga, Michael Neff, Christian Theobalt
摘要
Recent deep learning-based approaches have shown promising results for synthesizing plausible 3D human gestures from speech input. However, these approaches typically offer limited freedom to incorporate user control. Furthermore, training such models in a supervised manner often does not capture the multi-modal nature of the data, particularly because the same audio input can produce different gesture outputs. To address these problems, we present an approach for generating controllable 3D gestures that combines the advantage of database matching and deep generative modeling.
Our method predicts 3D body motion by sequentially searching for the most plausible audio-gesture clips from a database using a k-Nearest Neighbors (k-NN) algorithm that considers the similarity to both the input audio and the previous body pose information. To further improve the synthesis quality, we propose a conditional Generative Adversarial Network (cGAN) model to provide a datadriven refinement to the k-NN result by comparing its plausibility against the ground truth audio-gesture pairs. Our novel approach enables direct and more varied control manipulation that is not possible with prior learning-based counterparts. Our experiments show that our proposed approach outperforms recent models on control-based synthesis tasks using high-level signals such as motion statistics while enabling flexible and effective user control for lower-level signals. 1
• Computing methodologies → Motion processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion SynthesisMathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 被引用 192 次
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion ModelsSimon Alexanderson, Rajmund Nagy, Jonas Beskow, Gustav Eje HenterSIGGRAPH 2023 · 被引用 191 次
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP LatentsTenglong Ao, Zeyi Zhang, Libin LiuSIGGRAPH 2023 · 被引用 151 次
- SINC: Spatial Composition of 3D Human Motions for Simultaneous Action GenerationNikos Athanasiou, Mathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 被引用 69 次
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao 等SIGGRAPH 2024 · 被引用 39 次
它引用的顶会 Paper5
- Learned motion matchingDaniel Holden, Oussama Kanoun, Maksym Perepichka, Tiberiu PopaSIGGRAPH 2020 · 被引用 146 次
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersJing Li, Di Kang, Wenjie Pei, Xuefei Zhe 等ICCV 2021 · 被引用 144 次
- Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and SynthesisGilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori 等ICCV 2019 · 被引用 114 次
- Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression LearningUttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, Dinesh ManochaACM MM 2021 · 被引用 94 次
- SGToolkit: An Interactive Gesture Authoring Toolkit for Embodied Conversational AgentsYoungwoo Yoon, Keunwoo Park, Minsu Jang, Jaehong Kim 等UIST 2021 · 被引用 20 次
相关 Paper
- Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture GenerationFengqi Liu, Hexiang Wang, Jingyu Gong, Ran Yi 等ACM MM 2024 · 被引用 2 次
- Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional ControlZunnan Xu, Yachao Zhang, Sicheng Yang, Ronghui Li 等AAAI 2024 · 被引用 20 次
- Passing a Non-verbal Turing Test: Evaluatina Gesture Animations Generated from SpeechManuel Rebol, Christian Gütl, Krzysztof PietroszekIEEE VR 2021 · 被引用 31 次
- Body2Hands: Learning To Infer 3D Hands From Conversational Gesture Body DynamicsEvonne Ng, Shiry Ginosar, Trevor Darrell, Hanbyul JooCVPR 2021
- Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language ModelsBohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding 等SIGGRAPH 2025 · 被引用 5 次
