SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
Ming Li, Xin Gu, Fan Chen, Xiaoying Xing, Longyin Wen, Chen Chen, Sijie Zhu
Abstract
Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signals caused by the mismatch between editing instructions and original-edited image pairs. Recent efforts attempt to improve editing models through generating higher-quality edited images, pre-training on recognition tasks, or introducing vision-language models (VLMs) but fail to resolve this fundamental issue. In this paper, we offer a novel solution by constructing more effective editing instructions for given image pairs. This includes rectifying the editing instructions to better align with the original-edited image pairs and using contrastive editing instructions to further enhance their effectiveness. Specifically, we find that editing models exhibit specific generation attributes at different inference steps, independent of the text. Based on these prior attributes, we define a unified guide for VLMs to rectify editing instructions. However, there are some challenging editing scenarios that cannot be resolved solely with rectified instructions. To this end, we further construct contrastive supervision signals with positive and negative instructions and introduce them into the model training using triplet loss, thereby further facilitating supervision effectiveness. Our method does not require the VLM modules or pre-training tasks used in previous work, offering a more direct and efficient way to provide better supervision signals, and providing a novel, simple, and effective solution for instruction-based image editing. Results on multiple benchmarks demonstrate that our method significantly outperforms existing approaches. Compared with previous SOTA SmartEdit, we achieve 9.19% improvements on the Real-Edit benchmark with 30x less training data and 13x smaller model size.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang et al.ICLR 2026 · 34 citations
- CamEdit: Continuous Camera Parameter Control for Photorealistic Image EditingXinran Qin, Zhixin Wang, Fan Li, Haoyu Chen et al.NeurIPS 2025 · 17 citations
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image EditingWei Chow, Linfeng Li, Lingdong Kong, Zefeng Li et al.CVPR 2026 · 14 citations
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image EditingYucheng Liao, Jiajun Liang, Kaiqian Cui, Baoquan Zhao et al.CVPR 2026 · 6 citations
- ViPO: Visual Preference Optimization at ScaleMing Li, Jie Wu, Jiaxing Cui, Xiaojie Li et al.ICLR 2026 · 2 citations
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- InsightEdit: Towards Better Instruction Following for Image EditingYingjing Xu, Jie Kong, Jiazhi Wang, Xiao Pan et al.CVPR 2025
- X2Edit: Revisiting Arbitrary-Instruction Image Editing Through Self-Constructed Data and Task-Aware Representation LearningJian Ma, Xujie Zhu, Zihao Pan, Qirong Peng et al.AAAI 2026 · 15 citations
- UIP2P: Unsupervised Instruction-Based Image Editing via Edit Reversibility ConstraintEnis Simsar, Alessio Tonioni, Yongqin Xian, Thomas Hofmann et al.ICCV 2025 · 1 citation
- HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP ModelsZhixiang Wei, Guangting Wang, Xiaoxiao Ma, Ke Mei et al.ICCV 2025 · 1 citation
- Multi-Reward as Condition for Instruction-based Image EditingXin Gu, Ming Li, Libo Zhang, Fan Chen et al.ICLR 2025
