MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
Yannan He, Garvita Tiwari, Xiaohan Zhang, Pankaj Bora, Tolga Birdal, Jan Eric Lenssen, Gerard Pons-Moll
摘要
A figure dances ballet elegantly.
The person is doing a dance twirl. The person is sweeping the floor.
We introduce MoLingo, a text-to-motion (T2M) model that generates realistic, lifelike human motion by denoising in a continuous latent space. Recent works perform latent space diffusion, either on the whole latent at once or auto-regressively over multiple latents. In this paper, we study how to make diffusion on continuous motion latents work best. We focus on two questions: (1) how to build a semantically aligned latent space so diffusion becomes more effective, and (2) how to best inject text conditioning so the motion follows the description closely. We propose a semantic-aligned motion encoder trained with frame-level text labels so that latents with similar text meaning stay close, which makes the latent space more diffusion-friendly. We also compare single-token conditioning with a multi-token cross-attention scheme and find that cross-attention gives better motion realism and text-motion alignment. With semantically aligned latents, auto-regressive generation, and cross-attention text conditioning, our model sets a new state-of-the-art in human motion generation on standard metrics and in a user study. We will release our code and models for further research and downstream usage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
相关 Paper
- LGTM: Local-to-Global Text-Driven Human Motion Diffusion ModelHaowen Sun, Ruikun Zheng, Haibin Huang, Chongyang Ma 等SIGGRAPH 2024 · 被引用 12 次
- Make-An-Animation: Large-Scale Text-conditional 3D Human Motion GenerationSamaneh Azadi, Akbar Shah, Thomas Hayes, Devi Parikh 等ICCV 2023 · 被引用 70 次
- EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent SpaceJianrong Zhang, Hehe Fan, Yi YangCVPR 2025
- AMD: Autoregressive Motion DiffusionBo Han, Hao Peng, Minjing Dong, Yi Ren 等AAAI 2024 · 被引用 30 次
- ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided AlignmentWanjiang Weng, Xiaofeng Tan, Junbo Wang, Guo-Sen Xie 等AAAI 2026 · 被引用 6 次
