HIVE: Harnessing Human Feedback for Instructional Visual Editing
Shu Zhang, Xinyi Yang, Yihao Feng, Can Qin, Chia-Chih Chen, Ning Yu, Zeyuan Chen, Huan Wang, Silvio Savarese, Stefano Ermon, Caiming Xiong, Ran Xu
Abstract
Remove the red arch Add a moon in the background Add a sweater for the duck Change the plant color to blue Figure 1. We show four groups of representative results. In each triplet, from left to right are: the original image, InstructPix2Pix [7] using our data (IP2P-Ours), and HIVE. We observe that HIVE leads to more acceptable results than the model without human feedback. For instance, in the left two examples, IP2P-Ours understands the editing instruction "remove" and "change to blue" individually, but fails to understand the corresponding objects. Human feedback resolves this ambiguity, as shown in other examples as well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers76
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- MagicDrive: Street View Generation with Diverse 3D Geometry ControlRuiyuan Gao, Kai Chen, Enze Xie, Lanqing Hong et al.ICLR 2024 · 248 citations
- GenArtist: Multimodal LLM as an Agent for Unified Image Generation and EditingZhenyu Wang, Aoxue Li, Zhenguo Li, Xihui LiuNeurIPS 2024 · 162 citations
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
- AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based PoliciesXixi Hu, Qiang Liu, Xingchao Liu, Bo LiuNeurIPS 2024 · 73 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Focus on Your Instruction: Fine-grained and Multi-instruction Image Editing by Attention ModulationQin Guo, Tianwei LinCVPR 2024
- HQ-Edit: A High-Quality Dataset for Instruction-based Image EditingMude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi et al.ICLR 2025
- Instruction-based Image Manipulation by Watching How Things MoveMingdeng Cao, Xuaner Zhang, Yinqiang Zheng, Zhihao XiaCVPR 2025
- InsightEdit: Towards Better Instruction Following for Image EditingYingjing Xu, Jie Kong, Jiazhi Wang, Xiao Pan et al.CVPR 2025
- InstructPix2Pix: Learning to Follow Image Editing InstructionsTim Brooks, Aleksander Holynski, Alexei A. EfrosCVPR 2023
