MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm
Ziyan Guo, Zeyu Hu, De Wen Soh, Na Zhao
Abstract
Human motion generation and editing are key components of computer vision. However, current approaches in this field tend to offer isolated solutions tailored to specific tasks, which can be inefficient and impractical for real-world applications. While some efforts have aimed to unify motion-related tasks, these methods simply use different modalities as conditions to guide motion generation. Consequently, they lack editing capabilities, fine-grained control, and fail to facilitate knowledge sharing across tasks. To address these limitations and provide a versatile, unified framework capable of handling both human motion generation and editing, we introduce a novel paradigm: Motion-Condition-Motion, which enables the unified formulation of diverse tasks with three concepts: source motion, condition, and target motion. Based on this paradigm, we propose a unified framework, MotionLab, which incorporates rectified flows to learn the mapping from source motion to target motion, guided by the specified conditions. In MotionLab, we introduce the 1) MotionFlow Transformer to enhance conditional generation and editing without task-specific modules; 2) Aligned Rotational Position Encoding to guarantee the time synchronization between source motion and target motion; 3) Task Specified Instruction Modulation; and 4) Motion Curriculum Learning for effective multi-task learning and knowledge sharing across tasks. Notably, our MotionLab demonstrates promising generalization capabilities and inference efficiency across multiple benchmarks for human motion. Our code and additional video results are available at: https://diouo.github.io/motionlab.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language ModelsHaidong Xu, Guangwei Xu, Zhedong Zheng, Xiatian Zhu et al.NeurIPS 2025 · 5 citations
- DanceEditor: Towards Iterative Editable Music-Driven Dance Generation with Open-Vocabulary DescriptionsHengyuan Zhang, Zhe Li, Xingqun Qi, Mengze Li et al.ICCV 2025 · 3 citations
- PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum LearningYingjie Xi, Jian Jun Zhang, Xiaosong YangACM MM 2025 · 1 citation
- Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow DistillationDivyanshu Daiya, Aniket BeraCVPR 2026
- Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose PredictionQiongjie Cui, Pan Zhou, Jingjing Chen, Na ZhaoCVPR 2026
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity PredictionZhengyuan Li, Kai Cheng, Anindita Ghosh, Uttaran Bhattacharya et al.CVPR 2025
- MultiAct: Long-Term 3D Human Motion Generation from Multiple Action LabelsTaeryung Lee, Gyeongsik Moon, Kyoung Mu LeeAAAI 2023 · 62 citations
- UniMotion: A Unified Motion Framework for Simulation, Prediction and PlanningNan Song, Junzhe Jiang, Jingyu Li, Xiatian Zhu et al.NeurIPS 2025 · 2 citations
- MotionEdit: Benchmarking and Learning Motion-Centric Image EditingYixin Wan, Lei Ke, Wenhao Yu, Kai-Wei Chang et al.CVPR 2026 · 7 citations
- MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion TransformerPenghui Liu, Jiangshan Wang, Yutong Shen, Shanhui Mo et al.AAAI 2026 · 2 citations
