MGPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
Mingshuang Luo, Ruibing Hou, Zhuo Li, Hong Chang, Zimo Liu, Yaowei Wang, Shiguang Shan
Abstract
This paper presents MGPT, an advanced ultimodal, ultitask framework for otion comprehension and generation. MGPT operates on three fundamental principles. The first focuses on creating a unified representation space for various motion-relevant modalities. We employ discrete vector quantization for multimodal conditional signals, such as text, music and motion/dance, enabling seamless integration into a large language model (LLM) with a single vocabulary. The second involves modeling motion generation directly in the raw motion space. This strategy circumvents the information loss associated with a discrete tokenizer, resulting in more detailed and comprehensive motion generation. Third, MGPT learns to model the connections and synergies among various motion-relevant tasks. Text, the most familiar and well-understood modality for LLMs, is utilized as a bridge to establish connections between different motion tasks, facilitating mutual reinforcement. To our knowledge, MGPT is the first model capable of comprehending and generating motions based on multiple signals. Extensive experiments highlight MGPT's superior performance across various motion-relevant tasks and its powerful zero-shot generalization capabilities for extremely challenging tasks. Project page: https://github.com/luomingshuang/M3GPT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- MotionGPT3: Human Motion as a Second ModalityBingfan Zhu, Biao Jiang, Sunyi Wang, Shixiang Tang et al.ICLR 2026 · 43 citations
- GENMO: A GENeralist Model for Human MOtionJiefeng Li, Jinkun Cao, Haotian Zhang, Davis Rempe et al.ICCV 2025 · 15 citations
- MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion ParadigmZiyan Guo, Zeyu Hu, De Wen Soh, Na ZhaoICCV 2025 · 10 citations
- HIS-GPT: Towards 3D Human-In-Scene Multimodal UnderstandingJiahe Zhao, Ruibing Hou, Zejie Tian, Hong Chang et al.ICCV 2025 · 6 citations
- Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language ModelsBohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding et al.SIGGRAPH 2025 · 5 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
Related papers
- MotionGPT: Finetuned LLMs Are General-Purpose Motion GeneratorsYaqi Zhang, Di Huang, Bin Liu, Shixiang Tang et al.AAAI 2024 · 174 citations
- MoConVQ: Unified Physics-Based Motion Control via Scalable Discrete RepresentationsHeyuan Yao, Zhenhua Song, Yuyang Zhou, Tenglong Ao et al.SIGGRAPH 2024 · 34 citations
- MotionMaster: Generalizable Text-Driven Motion Generation and EditingNan Jiang, Yunhao Li, Lexi Pang, Zimo He et al.CVPR 2026
- ReMoGPT: Part-Level Retrieval-Augmented Motion-Language ModelsQing Yu, Mikihiro Tanaka, Kent FujiwaraAAAI 2025 · 6 citations
- A Motion is Worth a Hybrid Sentence: Taming Language Model for Unified Motion Generation by Fine-grained PlanningRonghui Li, Lingxiao Han, Shi Shu, Yueyao Liu et al.ACM MM 2025
