AMD: Autoregressive Motion Diffusion
Bo Han, Hao Peng, Minjing Dong, Yi Ren, Yixuan Shen, Chang Xu
摘要
Human motion generation aims to produce plausible human motion sequences according to various conditional inputs, such as text or audio. Despite the feasibility of existing methods in generating motion based on short prompts and simple motion patterns, they encounter difficulties when dealing with long prompts or complex motions. The challenges are two-fold: 1) the scarcity of human motion-captured data for long prompts and complex motions. 2) the high diversity of human motions in the temporal domain and the substantial divergence of distributions from conditional modalities, leading to a many-to-many mapping problem when generating motion with complex and long texts. In this work, we address these gaps by 1) elaborating the first dataset pairing long textual descriptions and 3D complex motions (HumanLong3D), and 2) proposing an autoregressive motion diffusion model (AMD). Specifically, AMD integrates the text prompt at the current timestep with the text prompt and action sequences at the previous timestep as conditional information to predict the current action sequences in an iterative manner. Furthermore, we present its generalization for X-to-Motion with “No Modality Left Behind”, enabling for the first time the generation of high-definition and high-fidelity human motions based on user-defined modality input.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Light-T2M: A Lightweight and Fast Model for Text-to-motion GenerationLing-An Zeng, Guohong Huang, Gaojie Wu, Wei-Shi ZhengAAAI 2025 · 被引用 23 次
- FlashMo: Geometric Interpolants and Frequency-Aware Sparsity for Scalable Efficient Motion GenerationZeyu Zhang, Yiran Wang, Danning Li, Dong Gong 等NeurIPS 2025 · 被引用 12 次
- RigAnything: Template-Free Autoregressive Rigging for Diverse 3D AssetsIsabella Liu, Zhan Xu, Wang Yifan, Hao Tan 等SIGGRAPH 2025 · 被引用 11 次
- Go to Zero: Towards Zero-Shot Motion Generation with Million-Scale DataKe Fan, Shunlin Lu, Minyue Dai, Runyi Yu 等ICCV 2025 · 被引用 11 次
- AnyTop: Character Animation Diffusion with Any TopologyInbar Gat, Sigal Raab, Guy Tevet, Yuval Reshef 等SIGGRAPH 2025 · 被引用 9 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
相关 Paper
- AMD: Anatomical Motion Diffusion with Interpretable Motion Decomposition and FusionBeibei Jing, Youjia Zhang, Zikai Song, Junqing Yu 等AAAI 2024 · 被引用 6 次
- Executing your Commands via Motion Diffusion in Latent SpaceXin Chen, Biao Jiang, Wen Liu, Zilong Huang 等CVPR 2023
- Synthesizing Long-Term Human Motions with Diffusion Models via Coherent SamplingZhao Yang, Bing Su, Ji-Rong WenACM MM 2023 · 被引用 17 次
- Flexible Motion In-betweening with Diffusion ModelsSetareh Cohan, Guy Tevet, Daniele Reda, Xue Bin Peng 等SIGGRAPH 2024 · 被引用 40 次
- Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene AffordanceZan Wang, Yixin Chen, Baoxiong Jia, Puhao Li 等CVPR 2024 · 被引用 38 次
