Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model
Yin Wang, Zhiying Leng, Frederick W. B. Li, Shun-Cheng Wu, Xiaohui Liang
Abstract
Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal and spatial relationships required to conform to a given text description. In this work, we propose a fine-grained method for generating high-quality, conditional human motion sequences supporting precise text description. Our approach consists of two key components: 1) a linguistics-structure assisted module that constructs accurate and complete language feature to fully utilize text information; and 2) a context-aware progressive reasoning module that learns neighborhood and overall semantic linguistics features from shallow and deep graph neural networks to achieve a multi-step inference. Experiments show that our approach outperforms text-driven motion generation methods on HumanML3D and KIT test sets and generates better visually confirmed motion to the text conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8a82f5b-69ee-43f4-a604-580afd4ead8eCited by top-tier papers46
- MoGenTS: Motion Generation based on Spatial-Temporal Joint ModelingWeihao Yuan, Yisheng He, Weichao Shen, Yuan Dong et al.NeurIPS 2024 · 51 citations
- MMM: Generative Masked Motion ModelEkkasit Pinyoanuntapong, Pu Wang, Minwoo Lee, Chen ChenCVPR 2024 · 39 citations
- Taming Diffusion Probabilistic Models for Character ControlRui Chen, Mingyi Shi, Shaoli Huang, Ping Tan et al.SIGGRAPH 2024 · 30 citations
- StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation FrameworkYiheng Huang, Hui Yang, Chuanchen Luo, Yuxi Wang et al.ACM MM 2024 · 27 citations
- Light-T2M: A Lightweight and Fast Model for Text-to-motion GenerationLing-An Zeng, Guohong Huang, Gaojie Wu, Wei-Shi ZhengAAAI 2025 · 23 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic GraphsPeng Jin, Yang Wu, Yanbo Fan, Zhongqian Sun et al.NeurIPS 2023 · 57 citations
- HGM³: Hierarchical Generative Masked Motion Modeling with Hard Token MiningMinjae Jeong, Yechan Hwang, Jaejin Lee, Sungyoon Jung et al.ICLR 2025
- AttT2M: Text-Driven Human Motion Generation with Multi-Perspective Attention MechanismChongyang Zhong, Lei Hu, Zihao Zhang, Shihong XiaICCV 2023 · 127 citations
- Modal-Enhanced Semantic Modeling for Fine-Grained 3D Human Motion RetrievalHaoyu Shi, Huaiwen ZhangACM MM 2024 · 3 citations
- Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene AffordanceZan Wang, Yixin Chen, Baoxiong Jia, Puhao Li et al.CVPR 2024 · 38 citations
