Motion-Aligned Word Embeddings for Text-to-Motion Generation
Ke Han, Yueming Lyu, Nicu Sebe
Abstract
Existing text-to-motion (T2M) generation models typically rely on pretrained large language models to encode textual inputs. However, these models, trained on generic text corpora, lack explicit alignment between motion-related words (e.g., "clockwise'', "quickly'') and human skeletal movements. This misalignment, fundamentally rooted in the word embedding layers, severely limits the ability of T2M models to understand and generalize fine-grained motion semantics. To tackle this issue, we propose Motion-Aligned Text Encoding (MATE), a novel framework that explicitly incorporates motion semantics into the word embedding layers of large language models to enhance text-motion alignment for motion generation. To address the challenge of inherent semantic entanglement in motion sequences, MATE introduces two key components: 1) a motion localization strategy that establishes localized correspondences between sub-texts and motion segments, enabling soft attention guidance for semantic localization; and 2) a motion disentanglement module that isolates word-specific motion semantics via contrastive kinematic prototypes, ensuring word-level alignment between linguistic and kinematic representations. Remarkably, language models enhanced with MATE can be seamlessly integrated into existing T2M methods, significantly surpassing state-of-the-art performance on two standard benchmarks with minimal modifications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7baec054-5241-47ec-b4ff-667cdf16eb4eBuilds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei et al.ICCV 2023 · 1,113 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
Related papers
- LGTM: Local-to-Global Text-Driven Human Motion Diffusion ModelHaowen Sun, Ruikun Zheng, Haibin Huang, Chongyang Ma et al.SIGGRAPH 2024 · 12 citations
- Zero-Shot Text-to-Motion Evaluation using Video Language ModelsYuwen Ji, Donglin Wang, Yue ZhangICML 2026
- Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time TrainingYunjiao Zhou, Xinyan Chen, Junlang Qian, Lihua Xie et al.AAAI 2026 · 2 citations
- Towards Unified Human Motion-Language Understanding via Sparse Interpretable CharacterizationGuangtao Lyu, Chenghao Xu, Jiexi Yan, Muli Yang et al.ICLR 2025
- FineXtrol: Controllable Motion Generation via Fine-Grained TextKeming Shen, Bizhu Wu, Junliang Chen, Xiaoqin Wang et al.AAAI 2026 · 3 citations
