Less is More: Improving Motion Diffusion Models with Sparse Keyframes
Jinseok Bae, Inwoo Hwang, Young Yoon Lee, Ziyu Guo, Joseph Liu, Yizhak Ben-Shabat, Young Min Kim, Mubbasir Kapadia
Abstract
Recent advances in motion diffusion models have led to remarkable progress in diverse motion generation tasks, including text-to-motion synthesis. However, existing approaches represent motions as dense frame sequences, requiring the model to process redundant or less informative frames. The processing of dense animation frames imposes significant training complexity, especially when learning intricate distributions of large motion datasets even with modern neural architectures. This severely limits the performance of generative motion models for downstream tasks. Inspired by professional animators who mainly focus on sparse keyframes, we propose a novel diffusion framework explicitly designed around sparse and geometrically meaningful keyframes. Our method reduces computation by masking non-keyframes and efficiently interpolating missing frames. We dynamically refine the keyframe mask during inference to prioritize informative frames in later diffusion steps. Extensive experiments show that our approach consistently outperforms state-of-the-art methods in text alignment and motion realism, while also effectively maintaining high performance at significantly fewer diffusion steps. We further validate the robustness of our framework by using it as a generative prior and adapting it to different downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54563f3e-56b2-484f-8270-4028a485946dCited by top-tier papers5
- FrankenMotion: Part-level Human Motion Generation and CompositionChuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger et al.CVPR 2026 · 10 citations
- Motion Synthesis with Sparse and Flexible Keyjoint ControlInwoo Hwang, Jinseok Bae, Donggeun Lim, Young Min KimICCV 2025 · 2 citations
- Scenemi: Motion In-Betweening for Modeling Human-Scene InteractionsInwoo Hwang, Bing Zhou, Young Min Kim, Jian Wang et al.ICCV 2025
- MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion GenerationWenfeng Song, Xuehan Wang, Shuai Li, Yi Chen et al.CVPR 2026
- EchoAvatar: Real-time Generative Avatar Animation from Audio StreamsBohong Chen, Yumeng Li, Yinglin Xu, Youyi Zheng et al.SIGGRAPH 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
Related papers
- Enhanced Fine-Grained Motion Diffusion for Text-Driven Human Motion SynthesisDong Wei, Xiaoning Sun, Huaijiang Sun, Shengxiang Hu et al.AAAI 2024 · 13 citations
- AutoKeyframe: Autoregressive Keyframe Generation for Human Motion Synthesis and EditingBowen Zheng, Ke Chen, Yuxin Yao, Zijiao Zeng et al.SIGGRAPH 2025 · 3 citations
- Towards Robust and Controllable Text-to-Motion via Masked Autoregressive DiffusionZongye Zhang, Bohan Kong, Qingjie Liu, Yunhong WangACM MM 2025 · 2 citations
- Flexible Motion In-betweening with Diffusion ModelsSetareh Cohan, Guy Tevet, Daniele Reda, Xue Bin Peng et al.SIGGRAPH 2024 · 40 citations
- Unifying Precise Keyframes and Semantic Control via Multi-level DiffusionLinjun Wu, Jiejia Yu, Leyang Jin, He Wang et al.CVPR 2026 · 1 citation
