MoMask: Generative Masked Modeling of 3D Human Motions
Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang, Li Cheng
2024Year
162Top-tier citations
Abstract
Walking forward and steps over an object, and then continue walking. Taking two strides forward, pivot swiftly on left foot, and then walk the other way. A person performs a standing back kick. Figure 1 . Our MoMask, when provided with a text input, generates high-quality 3D human motion with diversity and precise control over subtleties such as "two strides forward", "pivot on left foot", and "pivot swiftly".
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07978897-ce8f-49f3-a0a0-5656de32c8c8Cited by top-tier papers162
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng et al.ICML 2026 · 104 citations
- GestureLSM: Latent Shortcut Based Co-Speech Gesture Generation with Spatial-Temporal ModelingPinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu et al.ICCV 2025 · 54 citations
- MoGenTS: Motion Generation based on Spatial-Temporal Joint ModelingWeihao Yuan, Yisheng He, Weichao Shen, Yuan Dong et al.NeurIPS 2024 · 51 citations
- SnapMoGen: Human Motion Generation from Expressive TextsChuan Guo, Inwoo Hwang, Jian Wang, Bing ZhouNeurIPS 2025 · 50 citations
- MotionGPT3: Human Motion as a Second ModalityBingfan Zhu, Biao Jiang, Sunyi Wang, Shixiang Tang et al.ICLR 2026 · 43 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
Related papers
- Seamless Human Motion Composition with Blended Positional EncodingsGermán Barquero, Sergio Escalera, Cristina PalmeroCVPR 2024
- MoFusion: A Framework for Denoising-Diffusion-Based Motion SynthesisRishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, Christian TheobaltCVPR 2023
- AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision RewardHaonan Han, Xiangzuo Wu, Huan Liao, Zunnan Xu et al.CVPR 2025
- ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation ModelShunlin Lu, Jingbo Wang, Zeyu Lu, Ling-Hao Chen et al.CVPR 2025
- ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction GenerationLing-An Zeng, Guohong Huang, Yi-Lin Wei, Shengbo Gu et al.CVPR 2025
