ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model
Shunlin Lu, Jingbo Wang, Zeyu Lu, Ling-Hao Chen, Wenxun Dai, Junting Dong, Zhiyang Dou, Bo Dai, Ruimao Zhang
2025Year
12Top-tier citations
Abstract
The person is preparing to kick a football with a shooting motion. This includes transferring weight from one leg to the other, swinging one leg back and then forward in a kicking gesture, while the arms balance the body."
Figure 1. The generation results of ScaMo-3B with a text input. Our model could deal with abstract sentences and long sentences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49c71c24-16f4-4f43-9a8f-e5e3ef1dbbe9Cited by top-tier papers12
- SnapMoGen: Human Motion Generation from Expressive TextsChuan Guo, Inwoo Hwang, Jian Wang, Bing ZhouNeurIPS 2025 · 50 citations
- 🎧MOSPA: Human Motion Generation Driven by Spatial AudioShuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan et al.NeurIPS 2025 · 13 citations
- Go to Zero: Towards Zero-Shot Motion Generation with Million-Scale DataKe Fan, Shunlin Lu, Minyue Dai, Runyi Yu et al.ICCV 2025 · 11 citations
- MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent SpaceLixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan et al.ICCV 2025 · 11 citations
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language ModelsHaidong Xu, Guangwei Xu, Zhedong Zheng, Xiatian Zhu et al.NeurIPS 2025 · 5 citations
Builds on36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- MoMask: Generative Masked Modeling of 3D Human MotionsChuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang et al.CVPR 2024
- PhysicsFC: Learning User-Controlled Skills for a Physics-Based Football Player ControllerMinsu Kim, Eunho Jung, Yoonsang LeeSIGGRAPH 2025 · 4 citations
- ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction GenerationLing-An Zeng, Guohong Huang, Yi-Lin Wei, Shengbo Gu et al.CVPR 2025
- AnySkill: Learning Open-Vocabulary Physical Skill for Interactive AgentsJieming Cui, Tengyu Liu, Nian Liu, Yaodong Yang et al.CVPR 2024
- MixerMDM: Learnable Composition of Human Motion Diffusion ModelsPablo Ruiz-Ponce, Germán Barquero, Cristina Palmero, Sergio Escalera et al.CVPR 2025
