Multi-Reward as Condition for Instruction-based Image Editing
Xin Gu, Ming Li, Libo Zhang, Fan Chen, Longyin Wen, Tiejian Luo, Sijie Zhu
摘要
High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable Diffusion, DALL-E) which are not trained for image editing. Accordingly, these datasets suffer from inaccurate instruction following, poor detail preserving, and generation artifacts. In this paper, we propose to address the training data quality issue with multi-perspective reward data instead of refining the ground-truth image quality. 1) we first design a quantitative metric system based on best-in-class LVLM (Large Vision Language Model), i.e., GPT-4o in our case, to evaluate the generation quality from 3 perspectives, namely, instruction following, detail preserving, and generation quality. For each perspective, we collected quantitative score in 0 ∼ 5 and text descriptive feedback on the specific failure points in ground-truth edited images, resulting in a high-quality editing reward dataset, i.e., RewardEdit20K. 2) We further proposed a novel training framework to seamlessly integrate the metric output, regarded as multi-reward, into editing models to learn from the imperfect training triplets. During training, the reward scores and text descriptions are encoded as embeddings and fed into both the latent space and the U-Net of the editing models as auxiliary conditions. During inference, we set these additional conditions to the highest score with no text description for failure points, to aim at the best generation outcome. 3) We also build a challenging evaluation benchmark with real-world images/photos and diverse editing instructions, named Real-Edit. Experiments indicate that our multi-reward conditioned model outperforms its no-reward counterpart on two popular editing pipelines, i.e., InsPix2Pix and SmartEdit. Code is released at https://github.com/bytedance/Multi-Reward-Editing .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image EditingKeming Wu, Sicong Jiang, Max Ku, Ping Nie 等ICLR 2026 · 被引用 60 次
- Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement LearningKaihang Pan, Yang Wu, Wendong Bu, Kai Shen 等NeurIPS 2025 · 被引用 11 次
- WiseEdit: Benchmarking Cognition- and Creativity-Informed Image EditingKaihang Pan, Weile Chen, Haiyi Qiu, Qifan Yu 等CVPR 2026 · 被引用 9 次
- Selftok-Zero: Reinforcement Learning for Visual Generation via Discrete and Autoregressive Visual TokensBohan Wang, Mingze Zhou, Zhongqi Yue, Wang Lin 等NeurIPS 2025 · 被引用 1 次
- ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing EvaluationSherry X. Chen, Yi Wei, Luowei Zhou, Suren KumarICCV 2025 · 被引用 1 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- InsViE-1M: Effective Instruction-Based Video Editing with Elaborate Dataset ConstructionYuhui Wu, Liyi Chen, Ruibin Li, Shihao Wang 等ICCV 2025 · 被引用 6 次
- InsightEdit: Towards Better Instruction Following for Image EditingYingjing Xu, Jie Kong, Jiazhi Wang, Xiao Pan 等CVPR 2025
- EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward ModelingXin Luo, Jiahao Wang, Chenyuan Wu, Shitao Xiao 等ICLR 2026 · 被引用 63 次
- AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaQifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan 等CVPR 2025
- SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image EditingMing Li, Xin Gu, Fan Chen, Xiaoying Xing 等ICCV 2025 · 被引用 2 次
