FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing
Mingyuan Zhang, Huirong Li, Zhongang Cai, Jiawei Ren, Lei Yang, Ziwei Liu
Abstract
Text-driven motion generation has achieved substantial progress with the emergence of diffusion models. However, existing methods still struggle to generate complex motion sequences that correspond to fine-grained descriptions, depicting detailed and accurate spatio-temporal actions. This lack of fine controllability limits the usage of motion generation to a larger audience. To tackle these challenges, we present FineMoGen, a diffusion-based motion generation and editing framework that can synthesize fine-grained motions, with spatial-temporal composition to the user instructions. Specifically, FineMoGen builds upon diffusion model with a novel transformer architecture dubbed Spatio-Temporal Mixture Attention (SAMI). SAMI optimizes the generation of the global attention template from two perspectives: 1) explicitly modeling the constraints of spatio-temporal composition; and 2) utilizing sparsely-activated mixture-of-experts to adaptively extract fine-grained features. To facilitate a large-scale study on this new fine-grained motion generation task, we contribute the HuMMan-MoGen dataset, which consists of 2,968 videos and 102,336 fine-grained spatio-temporal descriptions. Extensive experiments validate that FineMoGen exhibits superior motion generation quality over state-of-the-art methods. Notably, FineMoGen further enables zero-shot motion editing capabilities with the aid of modern large language models (LLM), which faithfully manipulates motion sequences with fine-grained instructions. Project Page: https://mingyuan-zhang.github.io/projects/FineMoGen.html
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc4aa2bf-776d-4560-8e39-dccf8e92d591Cited by top-tier papers43
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 78 citations
- MoGenTS: Motion Generation based on Spatial-Temporal Joint ModelingWeihao Yuan, Yisheng He, Weichao Shen, Yuan Dong et al.NeurIPS 2024 · 51 citations
- The Quest for Generalizable Motion Generation: Data, Model, and EvaluationJing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang et al.ICLR 2026 · 23 citations
- Iterative Motion Editing with Natural LanguagePurvi Goel, Kuan-Chieh Wang, C. Karen Liu, Kayvon FatahalianSIGGRAPH 2024 · 22 citations
- CoMA: Compositional Human Motion Generation with Multi-modal AgentsShanlin Sun, Jiaqi Xu, Gabriel de Araujo, Shenghan Zhou et al.AAAI 2026 · 16 citations
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 371 citations
- FLAME: Free-Form Language-Based Motion Synthesis & EditingJihoon Kim, Jiseob Kim, Sungjoon ChoiAAAI 2023 · 276 citations
Related papers
- FineMotion: A Dataset and Benchmark with Both Spatial and Temporal Annotation for Fine-Grained Motion Generation and EditingBizhu Wu, Jinheng Xie, Meidan Ding, Zhe Kong et al.ICCV 2025 · 11 citations
- OMG: Towards Open-vocabulary Motion Generation via Mixture of ControllersHan Liang, Jiacheng Bao, Ruichi Zhang, Sihan Ren et al.CVPR 2024
- FrankenMotion: Part-level Human Motion Generation and CompositionChuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger et al.CVPR 2026 · 10 citations
- EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent SpaceJianrong Zhang, Hehe Fan, Yi YangCVPR 2025
- Modular-Cam: Modular Dynamic Camera-view Video Generation with LLMZirui Pan, Xin Wang, Yipeng Zhang, Hong Chen et al.AAAI 2025 · 6 citations
