Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
Bohong Chen, Yumeng Li, Yao-Xiang Ding, Tianjia Shao, Kun Zhou
Abstract
Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts, such as talking while walking. The major challenges lie in 1) the existing speech-to-motion datasets only involve highly limited full-body motions, making a wide range of common human activities out of training distribution; 2) these datasets also lack annotated user prompts. To address these challenges, we propose SynTalker, which utilizes the off-the-shelf text-to-motion dataset as an auxiliary for supplementing the missing full-body motion and prompts. The core technical contributions are two-fold. One is the multi-stage training process which obtains an aligned embedding space of motion, speech, and prompts despite the significant distributional mismatch in motion between speech-to-motion and text-to-motion datasets. Another is the diffusion-based conditional inference process, which utilizes the separate-then-combine strategy to realize fine-grained control of local body parts. Extensive experiments are conducted to verify that our approach supports precise and flexible control of synergistic full-body motion generation based on both speeches and user prompts, which is beyond the ability of existing approaches. Our code, pre-trained models, and videos are available at https://bohongchen.github.io/SynTalker-Page/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 093c6596-79e5-454a-92ed-c3b8a9a5f98eCited by top-tier papers12
- GestureLSM: Latent Shortcut Based Co-Speech Gesture Generation with Spatial-Temporal ModelingPinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu et al.ICCV 2025 · 54 citations
- SemTalk: Holistic Co-Speech Motion Generation with Frame-Level Semantic EmphasisXiangyue Zhang, Jianfang Li, Jiaxu Zhang, Ziqiang Dang et al.ICCV 2025 · 12 citations
- EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion GenerationXiangyue Zhang, Jianfang Li, Jiaxu Zhang, Jianqiang Ren et al.ACM MM 2025 · 7 citations
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual BodyJuze Zhang, Changan Chen, Xin Chen, Heng Yu et al.CVPR 2026 · 7 citations
- PyraMotion: Attentional Pyramid-Structured Motion Integration for Co-Speech 3D Gesture SynthesisZhizhuo Yin, Yuk Hang Tsui, Pan HuiNeurIPS 2025 · 6 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal ControlsYuxuan Bian, Ailing Zeng, Xuan Ju, Xian Liu et al.AAAI 2025 · 22 citations
- CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture GenerationFengyi Fang, Sicheng Yang, Wenming YangCVPR 2026 · 4 citations
- Towards High-fidelity 3D Talking Avatar with Personalized Dynamic TextureXuanchen Li, Jianyu Wang, Yuhao Cheng, Yikun Zeng et al.CVPR 2025
- Democratizing High-Fidelity Co-Speech Gesture Video GenerationXu Yang, Shaoli Huang, Shenbo Xie, Xuelin Chen et al.ICCV 2025 · 1 citation
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video GenerationShuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang et al.CVPR 2026 · 9 citations
