MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation
Yanbo Ding, Xirui Hu, Guo Zhi, Yan Zhang, Xinrui Wang, Zhixiang He, Chi Zhang, Yali Wang, Xuelong Li
Abstract
Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits generalization and discards essential 4D information for open-world animation. To address this, we propose MTVCraft (Motion Tokenization Video Crafter), the first framework that directly models raw 3D motion sequences (i.e., 4D motion) for character image animation. Specifically, we introduce 4DMoT (4D motion tokenizer) to quantize 3D motion sequences into 4D motion tokens. Compared to 2D-rendered pose images, 4D motion tokens offer more robust spatial-temporal cues and avoid strict pixel-level alignment between pose images and the character, enabling more flexible and disentangled control. Next, we introduce MV-DiT (Motion-aware Video DiT). By designing unique motion attention with 4D positional encodings, MV-DiT can effectively leverage motion tokens as 4D compact yet expressive context for character image animation in the complex 4D world. We implement MTVCraft on both CogVideoX-5B (small scale) and Wan-2.1-14B (large scale), demonstrating that our framework is easily scalable and can be applied to models of varying sizes. Experiments on the TikTok and Fashion benchmarks demonstrate our state-ofthe-art performance. Moreover, powered by robust motion tokens, MTVCraft showcases unparalleled zero-shot generalization. It can animate arbitrary characters in full-body and half-body forms, and even non-human objects across diverse styles and scenarios. Hence, it marks a significant step forward in this field and opens a new direction for pose-guided video generation. Our project page is available at https://github.com/DINGYANB/MTVCrafter . A scaled version has been commercially deployed and is available at https:// telestudio.teleagi.cn/generatevideo/creativeWorkshop .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7baa0d3d-c31d-4213-9d1d-455b9f26d335Cited by top-tier papers3
- 3D-Aware Implicit Motion Control for View-Adaptive Human Video GenerationZhixue Fang, Xu He, Songlin Tang, Haoxian Zhang et al.CVPR 2026 · 4 citations
- Superman: Unifying Skeleton and Vision for Human Motion Perception and GenerationXinshun Wang, Peiming Li, Ziyi Wang, Zhongbin Fang et al.CVPR 2026
- DreamActor-M2: Universal Character Image Animation via Spatiotemporal In-Context LearningMingshuang Luo, Shuang Liang, Zhengkun Rong, Yuxuan Luo et al.SIGGRAPH 2026
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- MultiAnimate: Pose-Guided Image Animation Made ExtensibleYingcheng Hu, Haowen Gong, Chuanguang Yang, Zhulin An et al.CVPR 2026 · 6 citations
- Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion ModelsMarc Benedí San Millán, Angela Dai, Matthias NießnerICLR 2026 · 3 citations
- MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image AnimationXirui Hu, Yanbo Ding, Jiahao Wang, Tingting Shi et al.ICLR 2026 · 3 citations
- Motion 3-to-4: 3D Motion Reconstruction for 4D SynthesisHongyuan Chen, Xingyu Chen, Zexiang Xu, Anpei ChenCVPR 2026 · 17 citations
- Motion4Motion: Motion Transfer Across Subjects at InferenceLing-Hao Chen, Zixin Yin, Duomin Wang, Xianfang Zeng et al.SIGGRAPH 2026
