Autoregressive Motion Generation with Gaussian Mixture-Guided Latent Sampling
Linnan Tu, Lingwei Meng, Zongyi Li, Hefei Ling, Shijuan Huang
摘要
Existing efforts in motion synthesis typically utilize either generative transformers with discrete representations or diffusion models with continuous representations. However, the discretization process in generative transformers can introduce motion errors, while the sampling process in diffusion models tends to be slow. In this paper, we propose a novel text-to-motion synthesis method GMMotion that combines a continuous motion representation with an autoregressive model, us-ing the Gaussian mixture model (GMM) to represent the conditional probability distribution. Unlike prior autoregressive approaches relying on residual vector quantization, our model employs continuous motion representations derived from the VAE’s latent space. This choice streamlines both the training and the inference processes while mitigating discretization errors. Specifically, we utilize a causal transformer to learn the distributions of continuous motion representations, which are modeled with a learnable Gaussian mixture model. Extensive experiments demonstrate that our model surpasses existing state-of-the-art models in the motion synthesis task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- BiMotion: B-spline Motion for Text-guided Dynamic 3D Character GenerationMiaowei Wang, Qingxuan Yan, Zhi Cao, Yayuan Li 等CVPR 2026 · 被引用 6 次
- MoLingo: Motion-Language Alignment for Text-to-Human Motion GenerationYannan He, Garvita Tiwari, Xiaohan Zhang, Pankaj Bora 等CVPR 2026 · 被引用 2 次
- TriC-Motion: Tri-Domain Causal Modeling Grounded Text-to-Motion GenerationYiyang Cao, Yunze Deng, Ziyu Lin, Bin Feng 等ICLR 2026
- MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention GuidanceNathan Sala, Ofir Abramovich, Ariel Shamir, Daniel Cohen-Or 等SIGGRAPH 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 被引用 672 次
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
相关 Paper
- Causal Motion Diffusion Models for Autoregressive Motion GenerationQing Yu, Akihisa Watanabe, Kent FujiwaraCVPR 2026 · 被引用 9 次
- MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent SpaceLixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan 等ICCV 2025 · 被引用 11 次
- Generating Human Motion from Textual Descriptions with Discrete RepresentationsJianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang 等CVPR 2023
- Towards Robust and Controllable Text-to-Motion via Masked Autoregressive DiffusionZongye Zhang, Bohan Kong, Qingjie Liu, Yunhong WangACM MM 2025 · 被引用 2 次
- Priority-Centric Human Motion Generation in Discrete Latent SpaceHanyang Kong, Kehong Gong, Dongze Lian, Michael Bi Mi 等ICCV 2023 · 被引用 81 次
