Scaling Large Motion Models with Million-Level Human Motions
Ye Wang, Sipeng Zheng, Bin Cao, Qianshan Wei, Weishuai Zeng, Qin Jin, Zongqing Lu
摘要
Inspired by the recent success of LLMs, the field of human motion understanding has increasingly shifted toward developing large motion models. Despite some progress, current efforts remain far from achieving truly generalist models, primarily due to the lack of massive high-quality data. To address this gap, we present MotionLib, the first million-level dataset for motion generation, which is at least 15× larger than existing counterparts and enriched with hierarchical text descriptions. Using MotionLib, we train a large motion model named Being-M0, demonstrating robust performance across a wide range of human activities, including unseen ones. Through systematic investigation, for the first time, we highlight the importance of scaling both data and model size for advancing motion generation, along with key insights to achieve this goal. To better integrate the motion modality, we propose Motionbook, an innovative motion encoding approach including (1) a compact yet lossless feature to represent motions; (2) a novel 2D lookup-free motion tokenizer that preserves fine-grained motion details while expanding codebook capacity, significantly enhancing the representational power of motion tokens. We believe this work lays the groundwork for developing more versatile and powerful motion generation models in the future. For further details, visit https://beingbeyond. github.io/Being-M0/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng 等ICML 2026 · 被引用 104 次
- CLUTCH: Contextualized Language model for Unlocking Text-Conditioned Hand motion modelling in the wildBalamurugan Thambiraja, Omid Taheri, Radek Danecek, Giorgio Becherini 等ICLR 2026 · 被引用 2 次
- RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion GenerationJiahao Zhang, Joseph Liu, Young-Yoon Lee, Seonghyeon Moon 等CVPR 2026 · 被引用 2 次
- Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid ControlWeisheng Xu, Qiwei Wu, Jiaxi Zhang, Jing Tan 等CVPR 2026 · 被引用 1 次
- FunPhase: A Periodic Functional Autoencoder for Motion Generation via Phase ManifoldsMarco Pegoraro, Evan Atherton, Bruno Roy, Aliasghar Khani 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Go to Zero: Towards Zero-Shot Motion Generation with Million-Scale DataKe Fan, Shunlin Lu, Minyue Dai, Runyi Yu 等ICCV 2025 · 被引用 11 次
- MotionMaster: Generalizable Text-Driven Motion Generation and EditingNan Jiang, Yunhao Li, Lexi Pang, Zimo He 等CVPR 2026
- MotionCtrl: A Real-Time Controllable Vision-Language-Motion ModelBin Cao, Sipeng Zheng, Ye Wang, Lujie Xia 等ICCV 2025 · 被引用 1 次
- SnapMoGen: Human Motion Generation from Expressive TextsChuan Guo, Inwoo Hwang, Jian Wang, Bing ZhouNeurIPS 2025 · 被引用 50 次
- OpenT2M: No-frill Motion Generation with Open-source, Large-scale, High-quality DataBin Cao, Sipeng Zheng, Hao Luo, Boyuan Li 等CVPR 2026 · 被引用 1 次
