EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space
Jianrong Zhang, Hehe Fan, Yi Yang
Abstract
Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diffusion models to effectively compose multiple semantic concepts into a single, coherent motion sequence. To address this issue, we propose EnergyMoGen, which includes two spectrums of Energy-Based Models: ❶ We interpret the diffusion model as a latent-aware energy-based model that generates motions by composing a set of diffusion models in latent space; ❷ We introduce a semanticaware energy model based on cross-attention, which enables semantic composition and adaptive gradient descent for text embeddings. To overcome the challenges of semantic inconsistency and motion distortion across these two spectrums, we introduce Synergistic Energy Fusion. This design allows the motion latent diffusion model to synthesize high-quality, complex motions by combining multiple energy terms corresponding to textual descriptions. Experiments show that our approach outperforms existing state-of-the-art models on various motion generation tasks, including text-to-motion generation, compositional motion generation, and multi-concept motion generation. Additionally, we demonstrate that our method can be used to extend motion datasets and improve the text-to-motion task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd01deda-e2dd-42a2-86ba-6ade7033291eCited by top-tier papers14
- Generative Trajectory Stitching through Diffusion CompositionYunhao Luo, Utkarsh A. Mishra, Yilun Du, Danfei XuNeurIPS 2025 · 48 citations
- FrankenMotion: Part-level Human Motion Generation and CompositionChuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger et al.CVPR 2026 · 10 citations
- Causal Motion Diffusion Models for Autoregressive Motion GenerationQing Yu, Akihisa Watanabe, Kent FujiwaraCVPR 2026 · 9 citations
- MotionHiFlow: Text-to-Motion via Hierarchical Flow MatchingHeng Li, Xiaotong Lin, Ling-An Zeng, Yulei Kang et al.CVPR 2026 · 7 citations
- ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided AlignmentWanjiang Weng, Xiaofeng Tan, Junbo Wang, Guo-Sen Xie et al.AAAI 2026 · 6 citations
Builds on43
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
Related papers
- Towards Decompositional Human Motion Generation with Energy-Based Diffusion ModelsJianrong Zhang, Hehe Fan, Yi YangCVPR 2026
- Event-T2M: Event-level Conditioning for Complex Text-to-Motion SynthesisSeong-Eun Hong, JaeYoung Seon, Juyeong Hwang, JongHwan Shin et al.ICLR 2026
- Energy-Based Cross Attention for Bayesian Context Update in Text-to-Image Diffusion ModelsGeon Yeong Park, Jeongsol Kim, Beomsu Kim, Sang Wan Lee et al.NeurIPS 2023 · 37 citations
- MoLingo: Motion-Language Alignment for Text-to-Human Motion GenerationYannan He, Garvita Tiwari, Xiaohan Zhang, Pankaj Bora et al.CVPR 2026 · 2 citations
- Hierarchical Enhancement of Semantic Priors for Disentangled Text-Driven Motion GenerationWenhan Lv, Shaopan Wang, Xiangyu Wu, Tianchu Hang et al.CVPR 2026
