CoMA: Compositional Human Motion Generation with Multi-modal Agents
Shanlin Sun, Jiaqi Xu, Gabriel de Araujo, Shenghan Zhou, Hanwen Zhang, Ziheng Huang, Chenyu You, Xiaohui Xie
摘要
3D human motion generation has seen substantial advancement in recent years. While state-of-the-art approaches have improved performance significantly, they still struggle with complex and detailed motions unseen in training data, largely due to the scarcity of motion datasets and the prohibitive cost of generating new training examples. To address these challenges, we introduce CoMA, an agent-based solution for complex human motion generation, editing, and comprehension. CoMA leverages multiple collaborative agents powered by large language and vision models, alongside a mask transformer-based motion generator featuring body part-specific encoders and codebooks for fine-grained control. Our framework enables generation of both short and long motion sequences with detailed instructions, text-guided motion editing, and self-correction for improved quality. Evaluations on the HumanML3D dataset demonstrate competitive performance against state-of-the-art methods. Additionally, we create a set of context-rich, compositional, and long text prompts, where user studies show our method significantly outperforms existing approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Generative Trajectory Stitching through Diffusion CompositionYunhao Luo, Utkarsh A. Mishra, Yilun Du, Danfei XuNeurIPS 2025 · 被引用 48 次
- HorizonForge: Driving Scene Editing with Any Trajectories and Any VehiclesYifan Wang, Francesco Pittaluga, Zaid Tasneem, Chenyu You 等CVPR 2026 · 被引用 3 次
- Let EEG Models Learn EEGYifan Wang, Yijia Ma, Wen Li, Chenyu YouICML 2026
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
相关 Paper
- Motion-Agent: A Conversational Framework for Human Motion Generation with LLMsQi Wu, Yubo Zhao, Yifan Wang, Xinhang Liu 等ICLR 2025
- MotionMaster: Generalizable Text-Driven Motion Generation and EditingNan Jiang, Yunhao Li, Lexi Pang, Zimo He 等CVPR 2026
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance DecouplingHaoyu Wang, Hao Tang, Donglin Di, Zhilu Zhang 等ICLR 2026 · 被引用 4 次
- MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple GranularitiesBizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong 等CVPR 2025
- ReMoGPT: Part-Level Retrieval-Augmented Motion-Language ModelsQing Yu, Mikihiro Tanaka, Kent FujiwaraAAAI 2025 · 被引用 6 次
