Autoregressive Motion Generation with Gaussian Mixture-Guided Latent Sampling
Linnan Tu, Lingwei Meng, Zongyi Li, Hefei Ling, Shijuan Huang
Abstract
Existing efforts in motion synthesis typically utilize either generative transformers with discrete representations or diffusion models with continuous representations. However, the discretization process in generative transformers can introduce motion errors, while the sampling process in diffusion models tends to be slow. In this paper, we propose a novel text-to-motion synthesis method GMMotion that combines a continuous motion representation with an autoregressive model, us-ing the Gaussian mixture model (GMM) to represent the conditional probability distribution. Unlike prior autoregressive approaches relying on residual vector quantization, our model employs continuous motion representations derived from the VAE’s latent space. This choice streamlines both the training and the inference processes while mitigating discretization errors. Specifically, we utilize a causal transformer to learn the distributions of continuous motion representations, which are modeled with a learnable Gaussian mixture model. Extensive experiments demonstrate that our model surpasses existing state-of-the-art models in the motion synthesis task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccbd4cbf-0b90-4a16-b693-9bf67f25dcd6Cited by top-tier papers4
- BiMotion: B-spline Motion for Text-guided Dynamic 3D Character GenerationMiaowei Wang, Qingxuan Yan, Zhi Cao, Yayuan Li et al.CVPR 2026 · 6 citations
- MoLingo: Motion-Language Alignment for Text-to-Human Motion GenerationYannan He, Garvita Tiwari, Xiaohan Zhang, Pankaj Bora et al.CVPR 2026 · 2 citations
- TriC-Motion: Tri-Domain Causal Modeling Grounded Text-to-Motion GenerationYiyang Cao, Yunze Deng, Ziyu Lin, Bin Feng et al.ICLR 2026
- MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention GuidanceNathan Sala, Ofir Abramovich, Ariel Shamir, Daniel Cohen-Or et al.SIGGRAPH 2026
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
Related papers
- Causal Motion Diffusion Models for Autoregressive Motion GenerationQing Yu, Akihisa Watanabe, Kent FujiwaraCVPR 2026 · 9 citations
- MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent SpaceLixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan et al.ICCV 2025 · 11 citations
- Generating Human Motion from Textual Descriptions with Discrete RepresentationsJianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang et al.CVPR 2023
- Towards Robust and Controllable Text-to-Motion via Masked Autoregressive DiffusionZongye Zhang, Bohan Kong, Qingjie Liu, Yunhong WangACM MM 2025 · 2 citations
- Priority-Centric Human Motion Generation in Discrete Latent SpaceHanyang Kong, Kehong Gong, Dongze Lian, Michael Bi Mi et al.ICCV 2023 · 81 citations
