Towards Emotion-enriched Text-to-Motion Generation via LLM-guided Limb-level Emotion Manipulating
Tan Yu, Jingjing Wang, Jiawen Wang, Jiamin Luo, Guodong Zhou
Abstract
In the literature, existing studies on text-to-motion generation (TMG) routinely focus on exploring the objective alignment of text and motion, which largely ignore the subjective emotion information, especially the limb-level emotion information. With this in mind, this paper proposes a new Emotion-enriched Text-to-Motion Generation (ETMG) task, aiming to generate motions with the subjective emotion information. Further, this paper believes that injecting emotions into limbs (named intra-limb emotion injection) and ensuring the coordination and coherence of emotional motions after injecting emotion information (named inter-limb emotion disturbance) is rather important and challenging in this ETMG task. To this end, this paper proposes an LL M-guided Limb-level Emotion Manipulating ( L3 EM) approach to ETMG. Specifically, this approach designs an LLM-guided intra-limb emotion modeling block to inject emotion into limbs, followed by a graph-structured inter-limb relation modeling block to ensure the coordination and coherence of emotional motions. Particularly, this paper constructs a coarse-grained Emotional Text-to-Motion (EmotionalT2M) dataset and a fine-grained Limb -level Emotional Text-to-Motion (Limb-ET2M) dataset to justify the effectiveness of the proposed L3EM approach. Detailed evaluation demonstrates the significant advantage of our L3EM approach to ETMG over the state-of-the-art baselines. This justifies the importance of the limb-level emotion information for ETMG and the effectiveness of our L3EM approach in coherently manipulating such information.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 79c48a9b-03b1-4cc9-8c5f-e7f0f194ff3bCited by top-tier papers2
- Tree-of-Reasoning: Towards Complex Medical Diagnosis via Multi-Agent Reasoning with Evidence TreeQi Peng, Jialin Cui, Jiayuan Xie, Yi Cai et al.ACM MM 2025 · 3 citations
- Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network for Multimodal Depression DetectionChangzeng Fu, Shiwen Zhao, Yunze Zhang, Zhongquan Jian et al.AAAI 2026
Related papers
- Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen ScenariosJiawen Wang, Jingjing Wang, Tianyang Chen, Min Zhang et al.AAAI 2026
- Towards LLM-centric Affective Visual Customization via Efficient and Precise Emotion ManipulatingJiamin Luo, Xuqian Gu, Jingjing Wang, Jiahong LuWWW 2026
- Modal-Enhanced Semantic Modeling for Fine-Grained 3D Human Motion RetrievalHaoyu Shi, Huaiwen ZhangACM MM 2024 · 3 citations
- LGTM: Local-to-Global Text-Driven Human Motion Diffusion ModelHaowen Sun, Ruikun Zheng, Haibin Huang, Chongyang Ma et al.SIGGRAPH 2024 · 12 citations
- MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple GranularitiesBizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong et al.CVPR 2025
