LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures
Chenghao Xu, Guangtao Lyu, Jiexi Yan, Muli Yang, Cheng Deng
Abstract
In response to the escalating demand for digital human representations, progress has been made in the generation of realistic human gestures from given speeches. Despite the remarkable achievements of recent research, the generation process frequently includes unintended, meaningless, or non-realistic gestures. To address this challenge, we propose a gesture translation paradigm, GesTran, which leverages large language models (LLMs) to deepen the understanding of the connection between speech and gesture and sequentially generates human gestures by interpreting gestures as a unique form of body language. The primary stage of the proposed framework employs a transformer-based auto-encoder network to encode human gestures into discrete symbols. Following this, the subsequent stage utilizes a pre-trained LLM to decipher the relationship between speech and gesture, translating the speech into gesture by interpreting the gesture as unique language tokens within the LLM. Our method has demonstrated state-of-the-art performance improvement through extensive and impartial experiments conducted on public TED and TED-Expressive datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3671c5f5-d485-4fef-9d17-969c5c86d2d5Cited by top-tier papers14
- Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang et al.ICCV 2025 · 2 citations
- Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative ModelingChenghao Xu, Guangtao Lyu, Jiexi Yan, Muli Yang et al.NeurIPS 2025 · 2 citations
- Enhancing Spoken Discourse Modeling in Language Models Using Gestural CuesVarsha Suresh, Muhammad Hamza Mughal, Christian Theobalt, Vera DembergACL 2025 · 2 citations
- VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual GroundingYihang Zhu, Jinhao Zhang, Yuxuan Wang, Aming Wu et al.ICCV 2025 · 1 citation
- ACFun: Abstract-Concrete Fusion Facial StylizationJiapeng Ji, Kun Wei, Ziqi Zhang, Cheng DengNeurIPS 2024 · 1 citation
Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- DanceFormer: Music Conditioned 3D Dance Generation with Parametric Motion TransformerBuyu Li, Yongchi Zhao, Zhelun Shi, Lu ShengAAAI 2022 · 182 citations
- MotionGPT: Finetuned LLMs Are General-Purpose Motion GeneratorsYaqi Zhang, Di Huang, Bin Liu, Shixiang Tang et al.AAAI 2024 · 174 citations
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP LatentsTenglong Ao, Zeyi Zhang, Libin LiuSIGGRAPH 2023 · 151 citations
Related papers
- Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language ModelsBohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding et al.SIGGRAPH 2025 · 5 citations
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao et al.SIGGRAPH 2024 · 39 citations
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with TransformerKunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost et al.SIGGRAPH 2023 · 18 citations
- Can Language Models Learn to Listen?Evonne Ng, Sanjay Subramanian, Dan Klein, Angjoo Kanazawa et al.ICCV 2023 · 44 citations
- The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human MotionChangan Chen, Juze Zhang, Shrinidhi K. Lakshmikanth, Yusu Fang et al.CVPR 2025
