Instruction-based Image Manipulation by Watching How Things Move
Mingdeng Cao, Xuaner Zhang, Yinqiang Zheng, Zhihao Xia
Abstract
2 Adobe Change the view to the side Lower the horse's head Make the man look angry Close the dog's eyes Have the man look at the side Move the camera to the left
Figure 1. We propose InstructMove, an instruction-based image editing model trained on frame pairs from videos with instructions generated by Multimodal LLMs. Our model excels at non-rigid editing, such as adjusting subject poses, expressions, and altering viewpoints, while maintaining content consistency. Additionally, our method supports precise, localized edits through the integration of masks, human poses, and other control mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e7f3515-beb7-4ced-b273-5e3021897c0bCited by top-tier papers6
- Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion TransformerZechuan Zhang, Ji Xie, Yu Lu, Zongxin Yang et al.NeurIPS 2025 · 18 citations
- Edicho: Consistent Image Editing in the WildQingyan Bai, Hao Ouyang, Yinghao Xu, Qiuyu Wang et al.ICCV 2025 · 10 citations
- PhotoFramer: Multi-modal Image Composition InstructionZhiyuan You, Ke Wang, He Zhang, Xin Cai et al.CVPR 2026 · 8 citations
- PICABench: How Far are We from Physical Realistic Image Editing?Yuandong Pu, Le Zhuo, Songhao Han, Jinbo Xing et al.ICLR 2026 · 7 citations
- Anchor Token Matching: Implicit Structure Locking for Training-Free AR Image EditingTaihang Hu, Linxuan Li, Kai Wang, Yaxing Wang et al.ICCV 2025 · 1 citation
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
- MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and EditingXueyun Tian, Wei Li, Bingbing Xu, Yige Yuan et al.ACM MM 2025 · 4 citations
- AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaQifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan et al.CVPR 2025
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image EditingYucheng Liao, Jiajun Liang, Kaiqian Cui, Baoquan Zhao et al.CVPR 2026 · 6 citations
- Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-TuningChenjian Gao, Lihe Ding, Xin Cai, Zhanpeng Huang et al.ICLR 2026 · 24 citations
