DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing
Chong Mou, Xintao Wang, Jiechong Song, Ying Shan, Jian Zhang
Abstract
Large-scale Text-to-Image (T2I) diffusion models have revolutionized image generation over the last few years. Although owning diverse and high-quality generation capabilities, translating these abilities to fine-grained Image editing remains challenging. In this paper, we propose DiffEditor to rectify two weaknesses in existing diffusion-based image editing: (1) in complex scenarios, editing results often lack editing accuracy and exhibit unexpected artifacts; (2) lack of flexibility to harmonize editing operations, e.g., imagine new content. In our solution, we introduce image prompts in fine-grained image editing, cooperating with the text prompt to better describe the editing content. To increase the flexibility while maintaining content consistency, we locally combine stochastic differential equation (SDE) into the ordinary differential equation (ODE) sampling. In addition, we incorporate regional score-based gradient guidance and a time travel strategy into the diffusion sampling, further improving the editing quality. Extensive experiments demonstrate that our method can efficiently achieve state-of-the-art performance on various fine-grained image editing tasks, including editing within a single image (e.g., object moving, resizing, and content dragging) and across images (e.g., appearance replacing and object pasting). Our source code is released at https://github.com/MC-E/DragonDiffusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb2854b6-bd79-480d-8db3-ce83e8ac3455Cited by top-tier papers47
- ReVideo: Remake a Video with Motion and Content ControlChong Mou, Mingdeng Cao, Xintao Wang, Zhaoyang Zhang et al.NeurIPS 2024 · 87 citations
- DragFlow: Unleashing DiT Priors with Region-Based Supervision for Drag EditingZihan Zhou, Shilin Lu, Shuli Leng, Shaocong Zhang et al.ICLR 2026 · 33 citations
- FastDrag: Manipulate Anything in One StepXuanjia Zhao, Jian Guan, Congyi Fan, Dongli Xu et al.NeurIPS 2024 · 27 citations
- RectifID: Personalizing Rectified Flow with Anchored Classifier GuidanceZhicheng Sun, Zhenhao Yang, Yang Jin, Haozhe Chi et al.NeurIPS 2024 · 13 citations
- Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention ReasonerXing Cui, Peipei Li, Zekun Li, Xuannan Liu et al.NeurIPS 2024 · 11 citations
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- DragonDiffusion: Enabling Drag-style Manipulation on Diffusion ModelsChong Mou, Xintao Wang, Jiechong Song, Ying Shan et al.ICLR 2024 · 223 citations
- PIXELS: Progressive Image Xemplar-based Editing with Latent SurgeryShristi Das Biswas, Matthew Shreve, Xuelu Li, Prateek Singhal et al.AAAI 2025 · 2 citations
- The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image EditingShen Nie, Hanzhong Allan Guo, Cheng Lu, Yuhao Zhou et al.ICLR 2024 · 64 citations
- FunEditor: Achieving Complex Image Edits via Function Aggregation with Diffusion ModelsMohammadreza Samadi, Fred X. Han, Mohammad Salameh, Hao Wu et al.AAAI 2025 · 1 citation
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 102 citations
