AttT2M: Text-Driven Human Motion Generation with Multi-Perspective Attention Mechanism
Chongyang Zhong, Lei Hu, Zihao Zhang, Shihong Xia
Abstract
Generating 3D human motion based on textual descriptions has been a research focus in recent years. It requires the generated motion to be diverse, natural, and conform to the textual description. Due to the complex spatio-temporal nature of human motion and the difficulty in learning the cross-modal relationship between text and motion, text-driven motion generation is still a challenging problem. To address these issues, we propose AttT2M, a two-stage method with multi-perspective attention mechanism: body-part attention and global-local motion-text attention. The former focuses on the motion embedding perspective, which means introducing a body-part spatio-temporal encoder into VQ-VAE to learn a more expressive discrete latent space. The latter is from the cross-modal perspective, which is used to learn the sentence-level and word-level motion-text cross-modal relationship. The text-driven motion is finally generated with a generative transformer. Extensive experiments conducted on HumanML3D and KIT-ML demonstrate that our method outperforms the current state-of-the-art works in terms of qualitative and quantitative evaluation, and achieve fine-grained synthesis and action2motion. Our code is in https://github.com/ZcyMonkey/AttT2M.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0201d691-9135-4e2c-b990-13d4d5d2c0ccCited by top-tier papers51
- MoGenTS: Motion Generation based on Spatial-Temporal Joint ModelingWeihao Yuan, Yisheng He, Weichao Shen, Yuan Dong et al.NeurIPS 2024 · 51 citations
- MMM: Generative Masked Motion ModelEkkasit Pinyoanuntapong, Pu Wang, Minwoo Lee, Chen ChenCVPR 2024 · 39 citations
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent GuidanceZhe Li, Yangyang Wei, Boan Zhu, Yibo Peng et al.ICLR 2026 · 29 citations
- Light-T2M: A Lightweight and Fast Model for Text-to-motion GenerationLing-An Zeng, Guohong Huang, Gaojie Wu, Wei-Shi ZhengAAAI 2025 · 23 citations
- LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive TokensZekun Li, Sizhe An, Chengcheng Tang, Chuan Guo et al.CVPR 2026 · 12 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Action2Motion: Conditioned Generation of 3D Human MotionsChuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou et al.ACM MM 2020 · 394 citations
Related papers
- Generating Human Motion from Textual Descriptions with Discrete RepresentationsJianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang et al.CVPR 2023
- Hierarchical Enhancement of Semantic Priors for Disentangled Text-Driven Motion GenerationWenhan Lv, Shaopan Wang, Xiangyu Wu, Tianchu Hang et al.CVPR 2026
- MotionHiFlow: Text-to-Motion via Hierarchical Flow MatchingHeng Li, Xiaotong Lin, Ling-An Zeng, Yulei Kang et al.CVPR 2026 · 7 citations
- Priority-Centric Human Motion Generation in Discrete Latent SpaceHanyang Kong, Kehong Gong, Dongze Lian, Michael Bi Mi et al.ICCV 2023 · 81 citations
- Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion ModelYin Wang, Zhiying Leng, Frederick W. B. Li, Shun-Cheng Wu et al.ICCV 2023 · 95 citations
