Video Motion Transfer with Diffusion Transformers
Alexander Pondaven, Aliaksandr Siarohin, Sergey Tulyakov, Philip Torr, Fabio Pizzati
Abstract
A lion walking through a busy market Robot walking on a sidewalk DiT Motorbiker driving around moonlit sand dunes Astronaut walking on the moon DiT Figure 1. Overview of DiTFlow. We propose a motion transfer method tailored for video Diffusion Transformers (DiT). We exploit a training-free strategy to transfer the motion of a reference video (top) to newly synthesized video content with arbitrary prompts (bottom). By optimizing DiT-specific positional embeddings, we can also synthesize new videos in a zero-shot manner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4539fd99-ba6b-47e1-ab6b-7d56b54c1d85Cited by top-tier papers17
- EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled FinetuningYue Ma, Yulong Liu, Qiyuan Zhu, Xiangpeng Yang et al.ICLR 2026 · 70 citations
- FastVMT: Eliminating Redundancy in Video Motion TransferYue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng et al.ICLR 2026 · 32 citations
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation ControlZekai Gu, Rui Yan, Jiahao Lu, Peng Li et al.SIGGRAPH 2025 · 21 citations
- RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion TransformersAhmet Berke Gökmen, Yigit Ekin, Bahri Batuhan Bilecen, Aysegul DundarNeurIPS 2025 · 14 citations
- Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock DenoisingAssaf Singer, Noam Rotstein, Amir Mann, Ron Kimmel et al.ICLR 2026 · 13 citations
Builds on35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- MotionFlow: Attention-Driven Motion Transfer in Video Diffusion ModelsTuna Han Salih Meral, Hidir Yesiltepe, Connor Dunlop, Pinar YanardagAAAI 2026
- Space-Time Diffusion Features for Zero-Shot Text-Driven Motion TransferDanah Yatim, Rafail Fridman, Omer Bar-Tal, Yoni Kasten et al.CVPR 2024 · 29 citations
- Object-WIPER: Training-Free Object and Associated Effect Removal in VideosSaksham Singh Kushwaha, Sayan Nag, Yapeng Tian, Kuldeep KulkarniCVPR 2026 · 5 citations
- Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward FeedbackXingpei Ma, Shenneng Huang, Jiaran Cai, Yuansheng Guan et al.AAAI 2026
- Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationChanggu Chen, Junwei Shu, Gaoqi He, Changbo Wang et al.AAAI 2025 · 1 citation
