LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures
Chenghao Xu, Guangtao Lyu, Jiexi Yan, Muli Yang, Cheng Deng
摘要
In response to the escalating demand for digital human representations, progress has been made in the generation of realistic human gestures from given speeches. Despite the remarkable achievements of recent research, the generation process frequently includes unintended, meaningless, or non-realistic gestures. To address this challenge, we propose a gesture translation paradigm, GesTran, which leverages large language models (LLMs) to deepen the understanding of the connection between speech and gesture and sequentially generates human gestures by interpreting gestures as a unique form of body language. The primary stage of the proposed framework employs a transformer-based auto-encoder network to encode human gestures into discrete symbols. Following this, the subsequent stage utilizes a pre-trained LLM to decipher the relationship between speech and gesture, translating the speech into gesture by interpreting the gesture as unique language tokens within the LLM. Our method has demonstrated state-of-the-art performance improvement through extensive and impartial experiments conducted on public TED and TED-Expressive datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang 等ICCV 2025 · 被引用 2 次
- Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative ModelingChenghao Xu, Guangtao Lyu, Jiexi Yan, Muli Yang 等NeurIPS 2025 · 被引用 2 次
- Enhancing Spoken Discourse Modeling in Language Models Using Gestural CuesVarsha Suresh, Muhammad Hamza Mughal, Christian Theobalt, Vera DembergACL 2025 · 被引用 2 次
- VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual GroundingYihang Zhu, Jinhao Zhang, Yuxuan Wang, Aming Wu 等ICCV 2025 · 被引用 1 次
- ACFun: Abstract-Concrete Fusion Facial StylizationJiapeng Ji, Kun Wei, Ziqi Zhang, Cheng DengNeurIPS 2024 · 被引用 1 次
它引用的顶会 Paper12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
- DanceFormer: Music Conditioned 3D Dance Generation with Parametric Motion TransformerBuyu Li, Yongchi Zhao, Zhelun Shi, Lu ShengAAAI 2022 · 被引用 182 次
- MotionGPT: Finetuned LLMs Are General-Purpose Motion GeneratorsYaqi Zhang, Di Huang, Bin Liu, Shixiang Tang 等AAAI 2024 · 被引用 174 次
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP LatentsTenglong Ao, Zeyi Zhang, Libin LiuSIGGRAPH 2023 · 被引用 151 次
相关 Paper
- Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language ModelsBohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding 等SIGGRAPH 2025 · 被引用 5 次
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao 等SIGGRAPH 2024 · 被引用 39 次
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with TransformerKunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost 等SIGGRAPH 2023 · 被引用 18 次
- Can Language Models Learn to Listen?Evonne Ng, Sanjay Subramanian, Dan Klein, Angjoo Kanazawa 等ICCV 2023 · 被引用 44 次
- The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human MotionChangan Chen, Juze Zhang, Shrinidhi K. Lakshmikanth, Yusu Fang 等CVPR 2025
