Hierarchical Enhancement of Semantic Priors for Disentangled Text-Driven Motion Generation
Wenhan Lv, Shaopan Wang, Xiangyu Wu, Tianchu Hang, Zhongquan Jian, Qingqiang Wu
Abstract
Text-to-motion generation aims to synthesize realistic and semantically aligned 3D human motions from natural language descriptions. Existing diffusion-based methods often rely on isotropic latent priors and shallow cross-modal supervision, which lead to semantic entanglement, limited controllability, and poor interpretability. We propose HESP, a unified diffusion framework that hierarchically enhances semantic priors for disentangled text-driven motion generation. At its core, HESP introduces an Adaptive Gaussian Variational Autoencoder (AG-VAE) that structures the latent motion manifold into multiple semantically coherent submanifolds, enabling interpretable and controllable motion representations. To further bridge linguistic and kinematic semantics, we design a Dynamic Cross-Modal Memory (DCMM) module for adaptive semantic fusion and a Hierarchical Cross-Modal Attention (HCA) mechanism to capture multi-level text-motion correspondences. Extensive experiments on HumanML3D and KIT-ML demonstrate that HESP consistently outperforms state-of-the-art baselines such as SALAD, MoMask, and MDM, achieving improvement while maintaining higher diversity and physical plausibility. Moreover, the structured latent space of HESP provides interpretable clusters that reveal clear semantic boundaries among different motion categories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f62e39c-5948-4f0e-b60c-4820fbaa2ac4Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- ReMoDiffuse: Retrieval-Augmented Motion Diffusion ModelMingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai et al.ICCV 2023 · 301 citations
Related papers
- Hierarchical Semantics Alignment for 3D Human Motion RetrievalYang Yang, Haoyu Shi, Huaiwen ZhangSIGIR 2024 · 4 citations
- Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic GraphsPeng Jin, Yang Wu, Yanbo Fan, Zhongqian Sun et al.NeurIPS 2023 · 57 citations
- MotionHiFlow: Text-to-Motion via Hierarchical Flow MatchingHeng Li, Xiaotong Lin, Ling-An Zeng, Yulei Kang et al.CVPR 2026 · 7 citations
- Towards Detailed Text-to-Motion Synthesis via Basic-to-Advanced Hierarchical Diffusion ModelZhenyu Xie, Yang Wu, Xuehao Gao, Zhongqian Sun et al.AAAI 2024 · 17 citations
- AttT2M: Text-Driven Human Motion Generation with Multi-Perspective Attention MechanismChongyang Zhong, Lei Hu, Zihao Zhang, Shihong XiaICCV 2023 · 127 citations
