MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
Yuxuan Bian, Ailing Zeng, Xuan Ju, Xian Liu, Zhaoyang Zhang, Wei Liu, Qiang Xu
摘要
Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to achieve various generation tasks with different condition modalities presents two main challenges: motion distribution drifts across different tasks (e.g., co-speech gestures and text-driven daily actions) and the complex optimization of mixed conditions with varying granularities (e.g., text and audio). Additionally, inconsistent motion formats across different tasks and datasets hinder effective training toward multimodal motion generation. In this paper, we propose MotionCraft, a unified diffusion transformer that crafts whole-body motion with plug-and-play multimodal control. Our framework employs a coarse-to-fine training strategy, starting with the first stage of text-to-motion semantic pre-training, followed by the second stage of multimodal low-level control adaptation to handle conditions of varying granularities. To effectively learn and transfer motion knowledge across different distributions, we design MC-Attn for parallel modeling of static and dynamic human topology graphs. To overcome the motion format inconsistency of existing benchmarks, we introduce MC-Bench, the first available multimodal whole-body motion generation benchmark based on the unified SMPL-X format. Extensive experiments show that MotionCraft achieves state-of-the-art performance on various standard motion generation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- The Quest for Generalizable Motion Generation: Data, Model, and EvaluationJing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang 等ICLR 2026 · 被引用 23 次
- MoAlign: Motion-Centric Representation Alignment for Video Diffusion ModelsAritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek, Amir Habibian 等ICLR 2026 · 被引用 15 次
- GENMO: A GENeralist Model for Human MOtionJiefeng Li, Jinkun Cao, Haotian Zhang, Davis Rempe 等ICCV 2025 · 被引用 15 次
- Video-As-Prompt: Unified Semantic Control for Video GenerationYuxuan Bian, Xin Chen, Zenan Li, Tiancheng Zhi 等ICLR 2026 · 被引用 13 次
- Do You Have Freestyle? Expressive Humanoid Locomotion via Audio ControlZhe Li, Cheng Chi, Yangyang Wei, Boan Zhu 等CVPR 2026 · 被引用 13 次
它引用的顶会 Paper25
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 被引用 536 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
- Action2Motion: Conditioned Generation of 3D Human MotionsChuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou 等ACM MM 2020 · 被引用 394 次
相关 Paper
- AMD: Autoregressive Motion DiffusionBo Han, Hao Peng, Minjing Dong, Yi Ren 等AAAI 2024 · 被引用 30 次
- Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion GenerationBohong Chen, Yumeng Li, Yao-Xiang Ding, Tianjia Shao 等ACM MM 2024 · 被引用 26 次
- Taming Diffusion Probabilistic Models for Character ControlRui Chen, Mingyi Shi, Shaoli Huang, Ping Tan 等SIGGRAPH 2024 · 被引用 30 次
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao 等ICML 2026 · 被引用 9 次
- StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross FusionZiyu Guo, Yizhak Ben-Shabat, Young Yoon Lee, Joseph Liu 等ICCV 2025 · 被引用 4 次
