ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
Fukun Yin, Shiyu Liu, Yucheng Han, Zhibo Wang, Peng Xing, Rui Wang, Wei Cheng, Yingming Wang, Aojie Li, Zixin Yin, Pengtao Chen, Xianfang Zeng
Abstract
Recent advances in image editing models have shown remarkable progress. A common architectural design couples a multimodal large language model (MLLM) encoder with a diffusion decoder, as seen in systems such as Step1X-Edit and Qwen-Image-Edit, where the MLLM encodes both the reference image and the instruction but remains frozen during training. In this work, we demonstrate that unlocking the reasoning capabilities of MLLM can further push the boundaries of editing models. Specifically, we explore two reasoning mechanisms, thinking and reflection, which enhance instruction understanding and editing accuracy. Based on that, our proposed framework enables image editing in a thinking-editing-reflection loop: the thinking mechanism leverages the world knowledge of MLLM to interpret abstract instructions, while the reflection reviews editing results, automatically corrects unintended manipulations, and identifies the stopping round. Extensive experiments demonstrate that our reasoning approach achieves significant performance gains, with improvements of ImgEdit (+4.3%), GEdit (+4.7%), and Kris (+8.2%) when initializing our DiT from the Step1X-Edit (ReasonEdit-S), and also outperforms previous open-source methods on both GEdit and Kris when integrated with Qwen-Image-Edit (ReasonEdit-Q).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f1233be-a84f-41bf-997c-939c4372f63cCited by top-tier papers1
Ask how each one uses itBuilds on25
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and ResolutionMostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek et al.NeurIPS 2023 · 303 citations
- OmniGen2: Towards Instruction-Aligned Multimodal GenerationChenyuan Wu, Jiahao Wang, Pengfei Zheng, Ruiran Yan et al.CVPR 2026 · 231 citations
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel et al.ICLR 2023 · 87 citations
Related papers
- Instruction-Based Image Editing with Planning, Reasoning, and GenerationLiya Ji, Chenyang Qi, Qifeng ChenICCV 2025 · 3 citations
- Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual ReasoningQingdong He, Xueqin Chen, Chaoyi Wang, Yanjie Pan et al.ICML 2026 · 6 citations
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
- InsightEdit: Towards Better Instruction Following for Image EditingYingjing Xu, Jie Kong, Jiazhi Wang, Xiao Pan et al.CVPR 2025
- Are Image-to-Video Models Good Zero-Shot Image Editors?Zechuan Zhang, Zhenyuan Chen, Zongxin Yang, Yi YangCVPR 2026 · 4 citations
