HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
Mude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi, Heng Wang, Peng Wang, Cihang Xie, Yuyin Zhou
摘要
a) Zoom in on the fox and add snowflakes falling around it. (b) Alter her hair color to black. (d) Replace the silver teapot with a ceramic blue and white patterned teapot with a similar wooden handle and a matching ceramic lid. (c) Replace the Amazon rainforest background with the underwater scenery of the Great Barrier Reef and adjust the parrot's position and wings to depict it flying. (e) Comparison between InstrictPix2Pix, HIVE, MagicBrush and HQ-Edit. Alignment Coherence Resolution HQ-Edit MagicBrush HIVE InstructPix2Pix Fig. 1: (a) -(d): example images and edit instructions from HQ-Edit. (e): we compare the dataset quality between our HQ-Edit and existing ones. Note that "Alignment" and "Coherence" are our newly developed metrics (introduced in Sec. 3.4) for measuring image/text qualities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper76
- I2EBench: A Comprehensive Benchmark for Instruction-based Image EditingYiwei Ma, Jiayi Ji, Ke Ye, Weihuang Lin 等NeurIPS 2024 · 被引用 67 次
- Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingYusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong 等CVPR 2026 · 被引用 63 次
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image EditingKeming Wu, Sicong Jiang, Max Ku, Ping Nie 等ICLR 2026 · 被引用 60 次
- EditVerse: Unifying Image and Video Editing and Generation with In-Context LearningXuan Ju, Tianyu Wang, Yuqian Zhou, He Zhang 等ICLR 2026 · 被引用 56 次
- UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow ModelsGuanlong Jiao, Biqing Huang, Kuan-Chieh Wang, Renjie LiaoICLR 2026 · 被引用 42 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji 等ICML 2024 · 被引用 786 次
相关 Paper
- HIVE: Harnessing Human Feedback for Instructional Visual EditingShu Zhang, Xinyi Yang, Yihao Feng, Can Qin 等CVPR 2024
- MoEdit: On Learning Quantity Perception for Multi-object Image EditingYanfeng Li, Ka-Hou Chan, Yue Sun, Chan-Tong Lam 等CVPR 2025
- Check, Locate, Rectify: A Training-Free Layout Calibration System for Text- to- Image GenerationBiao Gong, Siteng Huang, Yutong Feng, Shiwei Zhang 等CVPR 2024
- Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion TransformerZechuan Zhang, Ji Xie, Yu Lu, Zongxin Yang 等NeurIPS 2025 · 被引用 18 次
- Focus on Your Instruction: Fine-grained and Multi-instruction Image Editing by Attention ModulationQin Guo, Tianwei LinCVPR 2024
