MotionGPT: Human Motion as a Foreign Language
Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, Tao Chen
Abstract
Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multimodal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to human language, often perceived as a form of body language. By fusing language data with large-scale motion models,motionlanguage pre-training that can enhance the performance of motion-related tasks becomes feasible. Driven by this insight, we propose MotionGPT, a unified, versatile, and user-friendly motion-language model to handle multiple motion-relevant tasks. Specifically, we employ the discrete vector quantization for human motion and transfer 3D motion into motion tokens, similar to the generation process of word tokens. Building upon this "motion vocabulary", we perform language modeling on both motion and text in a unified manner, treating human motion as a specific language. Moreover, inspired by prompt learning, we pre-train MotionGPT with a mixture of motion-language data and fine-tune it on prompt-based questionand-answer tasks. Extensive experiments demonstrate that MotionGPT achieves state-of-the-art performances on multiple motion tasks including text-driven motion generation, motion captioning, motion prediction, and motion in-between. * Contributed equally and work done while Biao Jiang was a Research Intern with Tencent PCG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c33bd1c6-f233-44e8-a4ea-5a431638fdb2Cited by top-tier papers222
- OmniControl: Control Any Joint at Any Time for Human Motion GenerationYiming Xie, Varun Jampani, Lei Zhong, Deqing Sun et al.ICLR 2024 · 228 citations
- MeshXL: Neural Coordinate Field for Generative 3D Foundation ModelsSijin Chen, Xin Chen, Anqi Pang, Xianfang Zeng et al.NeurIPS 2024 · 125 citations
- Universal Humanoid Motion Representations for Physics-Based ControlZhengyi Luo, Jinkun Cao, Josh Merel, Alexander Winkler et al.ICLR 2024 · 125 citations
- HumanTOMATO: Text-aligned Whole-body Motion GenerationShunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin et al.ICML 2024 · 124 citations
- KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic SkillsWeiji Xie, Jinrui Han, Jiakun Zheng, Huanyu Li et al.NeurIPS 2025 · 120 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- ReMoGPT: Part-Level Retrieval-Augmented Motion-Language ModelsQing Yu, Mikihiro Tanaka, Kent FujiwaraAAAI 2025 · 6 citations
- MotionGPT3: Human Motion as a Second ModalityBingfan Zhu, Biao Jiang, Sunyi Wang, Shixiang Tang et al.ICLR 2026 · 43 citations
- MGPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and GenerationMingshuang Luo, Ruibing Hou, Zhuo Li, Hong Chang et al.NeurIPS 2024
- MotionGPT: Finetuned LLMs Are General-Purpose Motion GeneratorsYaqi Zhang, Di Huang, Bin Liu, Shixiang Tang et al.AAAI 2024 · 174 citations
- MotionMaster: Generalizable Text-Driven Motion Generation and EditingNan Jiang, Yunhao Li, Lexi Pang, Zimo He et al.CVPR 2026
