Priority-Centric Human Motion Generation in Discrete Latent Space
Hanyang Kong, Kehong Gong, Dongze Lian, Michael Bi Mi, Xinchao Wang
Abstract
Text-to-motion generation is a formidable task, aiming to produce human motions that align with the input text while also adhering to human capabilities and physical laws. While there have been advancements in diffusion models, their application in discrete spaces remains underexplored. Current methods often overlook the varying significance of different motions, treating them uniformly. It is essential to recognize that not all motions hold the same relevance to a particular textual description. Some motions, being more salient and informative, should be given precedence during generation. In response, we introduce a Priority-Centric Motion Discrete Diffusion Model (M2DM), which utilizes a Transformer-based VQ-VAE to derive a concise, discrete motion representation, incorporating a global self-attention mechanism and a regularization term to counteract code collapse. We also present a motion discrete diffusion model that employs an innovative noise schedule, determined by the significance of each motion token within the entire motion sequence. This approach retains the most salient motions during the reverse diffusion process, leading to more semantically rich and varied motions. Additionally, we formulate two strategies to gauge the importance of motion tokens, drawing from both textual and visual indicators. Comprehensive experiments on the HumanML3D and KIT-ML datasets confirm that our model surpasses existing techniques in fidelity and diversity, particularly for intricate textual descriptions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15d8364e-eae3-447d-ae23-b6ee16d4f396Cited by top-tier papers37
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 78 citations
- MoGenTS: Motion Generation based on Spatial-Temporal Joint ModelingWeihao Yuan, Yisheng He, Weichao Shen, Yuan Dong et al.NeurIPS 2024 · 51 citations
- MMM: Generative Masked Motion ModelEkkasit Pinyoanuntapong, Pu Wang, Minwoo Lee, Chen ChenCVPR 2024 · 39 citations
- StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation FrameworkYiheng Huang, Hui Yang, Chuanchen Luo, Yuxi Wang et al.ACM MM 2024 · 27 citations
- Light-T2M: A Lightweight and Fast Model for Text-to-motion GenerationLing-An Zeng, Guohong Huang, Gaojie Wu, Wei-Shi ZhengAAAI 2025 · 23 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- Generating Human Motion from Textual Descriptions with Discrete RepresentationsJianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang et al.CVPR 2023
- AttT2M: Text-Driven Human Motion Generation with Multi-Perspective Attention MechanismChongyang Zhong, Lei Hu, Zihao Zhang, Shihong XiaICCV 2023 · 127 citations
- Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked AutoregressionZichong Meng, Yiming Xie, Xiaogang Peng, Zeyu Han et al.CVPR 2025
- Causal Motion Diffusion Models for Autoregressive Motion GenerationQing Yu, Akihisa Watanabe, Kent FujiwaraCVPR 2026 · 9 citations
- GenM3: Generative Pretrained Multi-Path Motion Model for Text Conditional Human Motion GenerationJunyu Shi, Lijiang Liu, Yong Sun, Zhiyuan Zhang et al.ICCV 2025 · 8 citations
