Breaking The Limits of Text-conditioned 3D Motion Synthesis with Elaborative Descriptions
Yijun Qian, Jack Urbanek, Alexander G. Hauptmann, Jungdam Won
Abstract
Given its wide applications, there is increasing focus on generating 3D human motions from textual descriptions. Differing from the majority of previous works, which regard actions as single entities and can only generate short sequences for simple motions, we propose EMS, an elaborative motion synthesis model conditioned on detailed natural language descriptions. It generates natural and smooth motion sequences for long and complicated actions by factorizing them into groups of atomic actions. Meanwhile, it understands atomic-action level attributes (e.g., motion direction, speed, and body parts) and enables users to generate sequences of unseen complex actions from unique sequences of known atomic actions with independent attribute settings and timings applied. We evaluate our method on the KIT Motion-Language and BABEL benchmarks, where it outperforms all previous state-of-the-art with noticeable margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd597c7e-63f4-409d-a77c-e71a1989566fCited by top-tier papers5
- GENMO: A GENeralist Model for Human MOtionJiefeng Li, Jinkun Cao, Haotian Zhang, Davis Rempe et al.ICCV 2025 · 15 citations
- InfiniDreamer: Arbitrarily Long Human Motion Generation Via Segment Score DistillationWenjie Zhuo, Fan Ma, Hehe FanICCV 2025 · 6 citations
- Forecasting 3D Scanpaths in Egocentric VideoFiona Ryan, Ishwarya Ananthabhotla, Yijun Qian, Judy Hoffman et al.CVPR 2026 · 1 citation
- Seamless Human Motion Composition with Blended Positional EncodingsGermán Barquero, Sergio Escalera, Cristina PalmeroCVPR 2024
- Generating Human Motion in 3D Scenes from Text DescriptionsZhi Cen, Huaijin Pi, Sida Peng, Zehong Shen et al.CVPR 2024
Builds on20
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
Related papers
- Synthesis of Compositional Animations from Textual DescriptionsAnindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt et al.ICCV 2021 · 226 citations
- AMD: Anatomical Motion Diffusion with Interpretable Motion Decomposition and FusionBeibei Jing, Youjia Zhang, Zikai Song, Junqing Yu et al.AAAI 2024 · 6 citations
- Open the Motion Door: Atomic Motion Decomposition and Recomposition for Open-Vocabulary Motion GenerationKe Fan, Jiangning Zhang, Ran Yi, Jingyu Gong et al.CVPR 2026
- Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic GraphsPeng Jin, Yang Wu, Yanbo Fan, Zhongqian Sun et al.NeurIPS 2023 · 57 citations
- Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion ModelYin Wang, Zhiying Leng, Frederick W. B. Li, Shun-Cheng Wu et al.ICCV 2023 · 95 citations
