MotionDiff: Training-Free Zero-Shot Interactive Motion Editing via Flow-Assisted Multi-View Diffusion
Yikun Ma, Yiqing Li, Jiawei Wu, Xing Luo, Zhi Jin
Abstract
1 Generative models have made remarkable advancements and are capable of producing high-quality content. However, performing controllable editing with generative models remains challenging, due to their inherent uncertainty in outputs. This challenge is particularly pronounced in motion editing, which involves the processing of spatial information. While some physics-based generative methods have attempted to implement motion editing, they typically operate on single-view images with simple motions, such as translation and dragging. These methods struggle to handle complex motions, such as, rotation and stretching, and ensure multi-view consistency, often necessitating resource-intensive retraining. To address these challenges, we propose MotionDiff, a training-free zero-shot diffusion method that leverages optical flow for complex motion editing among multi-view images. Specifically, given a static scene, users can interactively select objects of interest to add motion priors. The proposed Point Kinematic Model (PKM) then estimates corresponding multi-view optical flows during the Multi-view Flow Estimation Stage (MFES). Subsequently, these optical flows are utilized to generate multiview motion results through decoupled motion representation in the Multi-view Motion Diffusion Stage (MMDS). Extensive experiments demonstrate that MotionDiff outperforms other physics-based generative motion editing methods in achieving high-quality multi-view consistent motion results. Notably, MotionDiff does not require retraining, enabling users to conveniently adapt it for various downstream tasks. Code is available at https://github.com/Mr-Ma-yikun/MotionDiff.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdbfc54e-3d04-4a27-b815-6a320f0e93a7Cited by top-tier papers1
Ask how each one uses itBuilds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Motion Guidance: Diffusion-Based Image Editing with Differentiable Motion EstimatorsDaniel Geng, Andrew OwensICLR 2024 · 46 citations
- COMD: Training-free Video Motion Transfer With Camera-Object Motion DisentanglementTeng Hu, Jiangning Zhang, Ran Yi, Yating Wang et al.ACM MM 2024 · 1 citation
- MotionCraft: Physics-Based Zero-Shot Video GenerationAntonio Montanaro, Luca Savant Aira, Emanuele Aiello, Diego Valsesia et al.NeurIPS 2024 · 52 citations
- MotionV2V: Editing Motion in a VideoRyan D. Burgert, Charles Herrmann, Forrester Cole, Michael S. Ryoo et al.CVPR 2026 · 13 citations
- MatchDiffusion: Training-Free Generation of Match-CutsAlejandro Pardo, Fabio Pizzati, Tong Zhang, Alexander Pondaven et al.ICCV 2025 · 3 citations
