Light-T2M: A Lightweight and Fast Model for Text-to-motion Generation
Ling-An Zeng, Guohong Huang, Gaojie Wu, Wei-Shi Zheng
Abstract
Despite the significant role text-to-motion (T2M) generation plays across various applications, current methods involve a large number of parameters and suffer from slow inference speeds, leading to high usage costs. To address this, we aim to design a lightweight model to reduce usage costs. First, unlike existing works that focus solely on global information modeling, we recognize the importance of local information modeling in the T2M task by reconsidering the intrinsic properties of human motion, leading us to propose a lightweight Local Information Modeling Module. Second, we introduce Mamba to the T2M task, reducing the number of parameters and GPU memory demands, and we have designed a novel Pseudo-bidirectional Scan to replicate the effects of a bidirectional scan without increasing parameter count. Moreover, we propose a novel Adaptive Textual Information Injector that more effectively integrates textual information into the motion during generation. By integrating the aforementioned designs, we propose a lightweight and fast model named Light-T2M. Compared to the state-of-the-art method, MoMask, our Light-T2M model features just 10% of the parameters (4.48M vs 44.85M) and achieves a 16% faster inference time (0.152s vs 0.180s), while surpassing MoMask with an FID of 0.040 (vs. 0.045) on HumanML3D dataset and 0.161 (vs. 0.228) on KIT-ML dataset. The code is available at https://github.com/qinghuannn/light-t2m .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f318571d-8702-4537-b750-43cb1201e4e8Cited by top-tier papers11
- Physics-Driven Spatiotemporal Modeling for AI-Generated Video DetectionShuhai Zhang, Zihao Lian, Jiahao Yang, Daiyuan Li et al.NeurIPS 2025 · 29 citations
- FlashMo: Geometric Interpolants and Frequency-Aware Sparsity for Scalable Efficient Motion GenerationZeyu Zhang, Yiran Wang, Danning Li, Dong Gong et al.NeurIPS 2025 · 12 citations
- MotionHiFlow: Text-to-Motion via Hierarchical Flow MatchingHeng Li, Xiaotong Lin, Ling-An Zeng, Yulei Kang et al.CVPR 2026 · 7 citations
- Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature RepresentationCongyi Fan, Jian Guan, Xuanjia Zhao, Dongli Xu et al.ICCV 2025 · 4 citations
- LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic InferenceJunkun JIANG, Ho Yin Au, Jingyu Xiang, Jie ChenCVPR 2026 · 2 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
Related papers
- MMM: Generative Masked Motion ModelEkkasit Pinyoanuntapong, Pu Wang, Minwoo Lee, Chen ChenCVPR 2024 · 39 citations
- Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense PredictionMang Cao, Sanping Zhou, Yizhe Li, Ye Deng et al.ICCV 2025
- LGTM: Local-to-Global Text-Driven Human Motion Diffusion ModelHaowen Sun, Ruikun Zheng, Haibin Huang, Chongyang Ma et al.SIGGRAPH 2024 · 12 citations
- MobileMamba: Lightweight Multi-Receptive Visual Mamba NetworkHaoyang He, Jiangning Zhang, Yuxuan Cai, Hongxu Chen et al.CVPR 2025
- Mamba-Reg: Vision Mamba Also Needs RegistersFeng Wang, Jiahao Wang, Sucheng Ren, Guoyizhe Wei et al.CVPR 2025
