MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
Yixin Wan, Lei Ke, Wenhao Yu, Kai-Wei Chang, Dong Yu
Abstract
We introduce MotionEdit , a novel dataset for motion-centric image editing—the task of modifying subject actions and interactions while preserving identity, structure, and physical plausibility.Unlike existing image editing datasets that focus on static appearance changes or contain only sparse, low-quality motion edits, MotionEdit provides high-fidelity image pairs depicting realistic motion transformations extracted and verified from continuous videos. This new task is not only scientifically challenging but also practically significant, powering downstream applications such as frame-controlled video synthesis and animation.To evaluate model performance on the novel task, we introduce MotionEdit-Bench , a benchmark that challenges models on motion-centric edits and measures model performance with generative, discriminative, and preference-based metrics.Benchmark results reveal that motion editing remains highly challenging for existing state-of-the-art diffusion-based editing models.To address this gap, we propose MotionNFT (Motion-guided Negative-aware FineTuning), a post-training framework that computes motion alignment rewards based on how well the motion flow between input and model-edited images matches the ground-truth motion, guiding models toward accurate motion transformations.Extensive experiments on FLUX.1 Kontext and Qwen-Image-Edit show that MotionNFT consistently improves editing quality and motion fidelity of both base models on the motion editing task without sacrificing general editing ability, demonstrating its effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 504d131a-3bfb-4699-a06d-2581b8a55472Cited by top-tier papers2
- CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement LearningYuhui WU, Chenxi Xie, Ruibin Li, Liyi Chen et al.ICML 2026 · 2 citations
- MotiMotion: Motion-Controlled Video Generation with Visual ReasoningHsin-Ying Lee, Hanwen Jiang, Yiqun Mei, Jing Shi et al.ICML 2026
Builds on11
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
- Learning to Estimate Hidden Motions with Global Motion AggregationShihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li et al.ICCV 2021 · 402 citations
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi et al.CVPR 2022 · 353 citations
Related papers
- X2Edit: Revisiting Arbitrary-Instruction Image Editing Through Self-Constructed Data and Task-Aware Representation LearningJian Ma, Xujie Zhu, Zihao Pan, Qirong Peng et al.AAAI 2026 · 15 citations
- Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion EditingGyojin Han, Junmo KimCVPR 2026 · 2 citations
- SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity PredictionZhengyuan Li, Kai Cheng, Anindita Ghosh, Uttaran Bhattacharya et al.CVPR 2025
- TexEditor: Structure-Preserving Text-Driven Texture EditingBo Zhao, Yihang Liu, Chenfeng Zhang, Huan Yang et al.ICML 2026 · 1 citation
- MotionV2V: Editing Motion in a VideoRyan D. Burgert, Charles Herrmann, Forrester Cole, Michael S. Ryoo et al.CVPR 2026 · 13 citations
