Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
Yusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong, Yinfei Yang, Jiasen Lu, Wenze Hu, Zhe Gan
摘要
Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community's progress remains constrained by the absence of large-scale, high-quality, and openly accessible datasets built from real images. We introduce Pico-Banana-400K, a comprehensive 400K-image dataset for instruction-based image editing. Our dataset is constructed by leveraging Nano-Banana to generate diverse edit pairs from real photographs in the OpenImages collection. What distinguishes Pico-Banana-400K from previous synthetic datasets is our systematic approach to quality and diversity. We employ a fine-grained image editing taxonomy to ensure comprehensive coverage of edit types while maintaining precise content preservation and instruction faithfulness through MLLM-based quality scoring and careful curation. Beyond single turn editing, Pico-Banana-400K enables research into complex editing scenarios. The dataset includes three specialized subsets: (1) a 72K-example multi-turn collection for studying sequential editing, reasoning, and planning across consecutive modifications; (2) a 56K-example preference subset for alignment research and reward model training; and (3) paired long-short editing instructions for developing instruction rewriting and summarization capabilities. By providing this large-scale, high-quality, and task-rich resource, Pico-Banana-400K establishes a robust foundation for training and benchmarking the next generation of text-guided image editing models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator OptimizationYunlong Lin, Linqing Wang, Kunjie Lin, Zixu Lin 等CVPR 2026 · 被引用 31 次
- ThinkGen: Generalized Thinking for Visual GenerationSiyu Jiao, Yiheng Lin, Yujie Zhong, Qi She 等CVPR 2026 · 被引用 12 次
- Are Image-to-Video Models Good Zero-Shot Image Editors?Zechuan Zhang, Zhenyuan Chen, Zongxin Yang, Yi YangCVPR 2026 · 被引用 4 次
- Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and AccessoriesJunyao Hu, Zhongwei Cheng, Waikeung Wong, Xingxing ZouCVPR 2026 · 被引用 4 次
- Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual GenerationJunjie Wang, 星华 娄, Xiangtai Li, Ye Tian 等ICML 2026
它引用的顶会 Paper17
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan 等ICCV 2023 · 被引用 770 次
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman 等ICLR 2023 · 被引用 361 次
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang 等ICLR 2024 · 被引用 173 次
- PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of CodeXuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu 等ICLR 2024 · 被引用 166 次
相关 Paper
- WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed BenchmarkWang Lin, Feng Wang, Majun Zhang, Wentao Hu 等ICLR 2026 · 被引用 2 次
- OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editingzhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu 等ICML 2026 · 被引用 24 次
- AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaQifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan 等CVPR 2025
- UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched EditsKeming Ye, Zhipeng Huang, Canmiao Fu, Qingyang Liu 等CVPR 2026 · 被引用 10 次
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image EditingKeming Wu, Sicong Jiang, Max Ku, Ping Nie 等ICLR 2026 · 被引用 60 次
