Lune

CVPR2025Top-tier venue

Instruction-based Image Manipulation by Watching How Things Move

Mingdeng Cao, Xuaner Zhang, Yinqiang Zheng, Zhihao Xia

2025Year
6Top-tier citations

Abstract

2 Adobe Change the view to the side Lower the horse's head Make the man look angry Close the dog's eyes Have the man look at the side Move the camera to the left

Figure 1. We propose InstructMove, an instruction-based image editing model trained on frame pairs from videos with instructions generated by Multimodal LLMs. Our model excels at non-rigid editing, such as adjusting subject poses, expressions, and altering viewpoints, while maintaining content consistency. Additionally, our method supports precise, localized edits through the integration of masks, human poses, and other control mechanisms.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 0e7f3515-beb7-4ced-b273-5e3021897c0b

Cited by top-tier papers6

Ask how each one uses it

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines