HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
Mude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi, Heng Wang, Peng Wang, Cihang Xie, Yuyin Zhou
Abstract
a) Zoom in on the fox and add snowflakes falling around it. (b) Alter her hair color to black. (d) Replace the silver teapot with a ceramic blue and white patterned teapot with a similar wooden handle and a matching ceramic lid. (c) Replace the Amazon rainforest background with the underwater scenery of the Great Barrier Reef and adjust the parrot's position and wings to depict it flying. (e) Comparison between InstrictPix2Pix, HIVE, MagicBrush and HQ-Edit. Alignment Coherence Resolution HQ-Edit MagicBrush HIVE InstructPix2Pix Fig. 1: (a) -(d): example images and edit instructions from HQ-Edit. (e): we compare the dataset quality between our HQ-Edit and existing ones. Note that "Alignment" and "Coherence" are our newly developed metrics (introduced in Sec. 3.4) for measuring image/text qualities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0cd87dd-378b-4af4-a9b7-d7088ad2e2ecCited by top-tier papers76
- I2EBench: A Comprehensive Benchmark for Instruction-based Image EditingYiwei Ma, Jiayi Ji, Ke Ye, Weihuang Lin et al.NeurIPS 2024 · 67 citations
- Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingYusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong et al.CVPR 2026 · 63 citations
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image EditingKeming Wu, Sicong Jiang, Max Ku, Ping Nie et al.ICLR 2026 · 60 citations
- EditVerse: Unifying Image and Video Editing and Generation with In-Context LearningXuan Ju, Tianyu Wang, Yuqian Zhou, He Zhang et al.ICLR 2026 · 56 citations
- UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow ModelsGuanlong Jiao, Biqing Huang, Kuan-Chieh Wang, Renjie LiaoICLR 2026 · 42 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
Related papers
- HIVE: Harnessing Human Feedback for Instructional Visual EditingShu Zhang, Xinyi Yang, Yihao Feng, Can Qin et al.CVPR 2024
- MoEdit: On Learning Quantity Perception for Multi-object Image EditingYanfeng Li, Ka-Hou Chan, Yue Sun, Chan-Tong Lam et al.CVPR 2025
- Check, Locate, Rectify: A Training-Free Layout Calibration System for Text- to- Image GenerationBiao Gong, Siteng Huang, Yutong Feng, Shiwei Zhang et al.CVPR 2024
- Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion TransformerZechuan Zhang, Ji Xie, Yu Lu, Zongxin Yang et al.NeurIPS 2025 · 18 citations
- Focus on Your Instruction: Fine-grained and Multi-instruction Image Editing by Attention ModulationQin Guo, Tianwei LinCVPR 2024
