A Motion Matching-based Framework for Controllable Gesture Synthesis from Speech
Ikhsanul Habibie, Mohamed A. Elgharib, Kripasindhu Sarkar, Ahsan Abdullah, Simbarashe Nyatsanga, Michael Neff, Christian Theobalt
Abstract
Recent deep learning-based approaches have shown promising results for synthesizing plausible 3D human gestures from speech input. However, these approaches typically offer limited freedom to incorporate user control. Furthermore, training such models in a supervised manner often does not capture the multi-modal nature of the data, particularly because the same audio input can produce different gesture outputs. To address these problems, we present an approach for generating controllable 3D gestures that combines the advantage of database matching and deep generative modeling.
Our method predicts 3D body motion by sequentially searching for the most plausible audio-gesture clips from a database using a k-Nearest Neighbors (k-NN) algorithm that considers the similarity to both the input audio and the previous body pose information. To further improve the synthesis quality, we propose a conditional Generative Adversarial Network (cGAN) model to provide a datadriven refinement to the k-NN result by comparing its plausibility against the ground truth audio-gesture pairs. Our novel approach enables direct and more varied control manipulation that is not possible with prior learning-based counterparts. Our experiments show that our proposed approach outperforms recent models on control-based synthesis tasks using high-level signals such as motion statistics while enabling flexible and effective user control for lower-level signals. 1
• Computing methodologies → Motion processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6368271-c2fb-4ca7-b1d3-3f7aedf6bcf6Cited by top-tier papers23
- TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion SynthesisMathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 192 citations
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion ModelsSimon Alexanderson, Rajmund Nagy, Jonas Beskow, Gustav Eje HenterSIGGRAPH 2023 · 191 citations
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP LatentsTenglong Ao, Zeyi Zhang, Libin LiuSIGGRAPH 2023 · 151 citations
- SINC: Spatial Composition of 3D Human Motions for Simultaneous Action GenerationNikos Athanasiou, Mathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 69 citations
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao et al.SIGGRAPH 2024 · 39 citations
Builds on5
- Learned motion matchingDaniel Holden, Oussama Kanoun, Maksym Perepichka, Tiberiu PopaSIGGRAPH 2020 · 146 citations
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersJing Li, Di Kang, Wenjie Pei, Xuefei Zhe et al.ICCV 2021 · 144 citations
- Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and SynthesisGilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori et al.ICCV 2019 · 114 citations
- Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression LearningUttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, Dinesh ManochaACM MM 2021 · 94 citations
- SGToolkit: An Interactive Gesture Authoring Toolkit for Embodied Conversational AgentsYoungwoo Yoon, Keunwoo Park, Minsu Jang, Jaehong Kim et al.UIST 2021 · 20 citations
Related papers
- Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture GenerationFengqi Liu, Hexiang Wang, Jingyu Gong, Ran Yi et al.ACM MM 2024 · 2 citations
- Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional ControlZunnan Xu, Yachao Zhang, Sicheng Yang, Ronghui Li et al.AAAI 2024 · 20 citations
- Passing a Non-verbal Turing Test: Evaluatina Gesture Animations Generated from SpeechManuel Rebol, Christian Gütl, Krzysztof PietroszekIEEE VR 2021 · 31 citations
- Body2Hands: Learning To Infer 3D Hands From Conversational Gesture Body DynamicsEvonne Ng, Shiry Ginosar, Trevor Darrell, Hanbyul JooCVPR 2021
- Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language ModelsBohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding et al.SIGGRAPH 2025 · 5 citations
