Synthesizing Long-Term Human Motions with Diffusion Models via Coherent Sampling
Zhao Yang, Bing Su, Ji-Rong Wen
摘要
Text-to-motion generation has gained increasing attention, but most existing methods are limited to generating short-term motions that correspond to a single sentence describing a single action. However, when a text stream describes a sequence of continuous motions, the generated motions corresponding to each sentence may not be coherently linked. Existing long-term motion generation methods face two main issues. Firstly, they cannot directly generate coherent motions and require additional operations such as interpolation to process the generated actions. Secondly, they generate subsequent actions in an autoregressive manner without considering the influence of future actions on previous ones. To address these issues, we propose a novel approach that utilizes a past-conditioned diffusion model with two optional coherent sampling methods: Past Inpainting Sampling and Compositional Transition Sampling. Past Inpainting Sampling completes subsequent motions by treating previous motions as conditions, while Compositional Transition Sampling models the distribution of the transition as the composition of two adjacent motions guided by different text prompts. Our experimental results demonstrate that our proposed method is capable of generating compositional and coherent long-term 3D human motions controlled by a user-instructed long text stream. The code is available at https://github.com/yangzhao1230/PCMDM
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- MMHead: Towards Fine-grained Multi-modal 3D Facial AnimationSijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan 等ACM MM 2024 · 被引用 17 次
- Enhancing Reward Models for High-Quality Image Generation: Beyond Text-Image AlignmentYing Ba, Tianyu Zhang, Yalong Bai, Wenyi Mo 等ICCV 2025 · 被引用 13 次
- MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent SpaceLixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan 等ICCV 2025 · 被引用 11 次
- Deep Compositional Phase Diffusion for Long Motion Sequence GenerationHo Yin Au, Jie Chen, Junkun Jiang, Jingyu XiangNeurIPS 2025 · 被引用 7 次
- InfiniDreamer: Arbitrarily Long Human Motion Generation Via Segment Score DistillationWenjie Zhuo, Fan Ma, Hehe FanICCV 2025 · 被引用 6 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- AMD: Autoregressive Motion DiffusionBo Han, Hao Peng, Minjing Dong, Yi Ren 等AAAI 2024 · 被引用 30 次
- Flexible Motion In-betweening with Diffusion ModelsSetareh Cohan, Guy Tevet, Daniele Reda, Xue Bin Peng 等SIGGRAPH 2024 · 被引用 40 次
- MultiAct: Long-Term 3D Human Motion Generation from Multiple Action LabelsTaeryung Lee, Gyeongsik Moon, Kyoung Mu LeeAAAI 2023 · 被引用 62 次
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 被引用 371 次
- DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion ControlKaifeng Zhao, Gen Li, Siyu TangICLR 2025 · 被引用 1 次
