Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
Zongye Zhang, Bohan Kong, Qingjie Liu, Yunhong Wang
摘要
Generating 3D human motion from text descriptions remains challenging due to the diverse and complex nature of human motion. While existing methods excel within the training distribution, they often struggle with out-of-distribution motions, limiting their applicability in real-world scenarios. Existing VQVAE-based methods often fail to represent novel motions faithfully using discrete tokens, which hampers their ability to generalize beyond seen data. Meanwhile, diffusion-based methods operating on continuous representations often lack fine-grained control over individual frames. To address these challenges, we propose a robust motion generation framework MoMADiff, which combines masked modeling with diffusion processes to generate motion using frame-level continuous representations. Our model supports flexible user-provided keyframe specification, enabling precise control over both spatial and temporal aspects of motion synthesis. MoMADiff demonstrates strong generalization capability on novel text-to-motion datasets with sparse keyframes as motion prompts. Extensive experiments on two held-out datasets and two standard benchmarks show that our method consistently outperforms state-of-the-art models in motion quality, instruction fidelity, and keyframe adherence. The code is available at: https://github.com/zzysteve/MoMADiff
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MotionHiFlow: Text-to-Motion via Hierarchical Flow MatchingHeng Li, Xiaotong Lin, Ling-An Zeng, Yulei Kang 等CVPR 2026 · 被引用 7 次
- Semantic-Aware Motion Encoding for Topology-Agnostic Character AnimationZongye Zhang, Yuzhuo Cui, Qingjie Liu, Yunhong WangICML 2026 · 被引用 1 次
- MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion GenerationWenfeng Song, Xuehan Wang, Shuai Li, Yi Chen 等CVPR 2026
- Zero-Shot Text-to-Motion Evaluation using Video Language ModelsYuwen Ji, Donglin Wang, Yue ZhangICML 2026
- TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-AnimationCheng-Feng Pu, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao 等SIGGRAPH 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked AutoregressionZichong Meng, Yiming Xie, Xiaogang Peng, Zeyu Han 等CVPR 2025
- ReMoDiffuse: Retrieval-Augmented Motion Diffusion ModelMingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai 等ICCV 2023 · 被引用 301 次
- InterMask: 3D Human Interaction Generation via Collaborative Masked ModelingMuhammad Gohar Javed, Chuan Guo, Li Cheng, Xingyu LiICLR 2025
- Hierarchical Enhancement of Semantic Priors for Disentangled Text-Driven Motion GenerationWenhan Lv, Shaopan Wang, Xiangyu Wu, Tianchu Hang 等CVPR 2026
- Less is More: Improving Motion Diffusion Models with Sparse KeyframesJinseok Bae, Inwoo Hwang, Young Yoon Lee, Ziyu Guo 等ICCV 2025 · 被引用 4 次
