Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control
Zunnan Xu, Yachao Zhang, Sicheng Yang, Ronghui Li, Xiu Li
Abstract
This study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated gestures. However, these methods perform poorly when certain modalities are missing during inference. To address this problem, we suggest using speech-derived multimodal priors to improve gesture generation. We introduce a novel method that separates priors from speech and employs multimodal priors as constraints for generating gestures. Our approach utilizes a chain-like modeling method to generate facial blendshapes, body movements, and hand gestures sequentially. Specifically, we incorporate rhythm cues derived from facial deformation and stylization prior based on speech emotions, into the process of generating gestures. By incorporating multimodal priors, our method improves the quality of generated gestures and eliminate the need for expensive setup preparation during inference. Extensive experiments and user studies confirm that our proposed approach achieves state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space ModelsZunnan Xu, Yukang Lin, Haonan Han, Sicheng Yang et al.NeurIPS 2024 · 62 citations
- GestureLSM: Latent Shortcut Based Co-Speech Gesture Generation with Spatial-Temporal ModelingPinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu et al.ICCV 2025 · 54 citations
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao et al.SIGGRAPH 2024 · 39 citations
- SemTalk: Holistic Co-Speech Motion Generation with Frame-Level Semantic EmphasisXiangyue Zhang, Jianfang Li, Jiaxu Zhang, Ziqiang Dang et al.ICCV 2025 · 12 citations
- CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture GenerationFengyi Fang, Sicheng Yang, Wenming YangCVPR 2026 · 4 citations
Builds on10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP LatentsTenglong Ao, Zeyi Zhang, Libin LiuSIGGRAPH 2023 · 151 citations
- Perturbed Self-Distillation: Weakly Supervised Large-Scale Point Cloud Semantic SegmentationYachao Zhang, Yanyun Qu, Yuan Xie, Zonghao Li et al.ICCV 2021 · 138 citations
Related papers
- Emotional Speech-Driven 3D Body Animation via Disentangled Latent DiffusionKiran Chhatre, Radek Danecek, Nikos Athanasiou, Giorgio Becherini et al.CVPR 2024
- SemGes: Semantics-Aware Co-Speech Gesture Generation Using Semantic Coherence and Relevance LearningLanmiao Liu, Esam Ghaleb, Asli Özyürek, Zerrin YumakICCV 2025 · 4 citations
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture ModelingHaiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng et al.CVPR 2024 · 55 citations
- SemGesture: Synthesizing Semantically Enhanced and Coherent GesturesPengsheng Liu, Zhaojie Chu, Xiaofen Xing, Xiangmin XuACM MM 2025 · 2 citations
- Audio-driven Neural Gesture Reenactment with Video Motion GraphsYang Zhou, Jimei Yang, Dingzeyu Li, Jun Saito et al.CVPR 2022 · 22 citations
