Constantly Improving Image Models Need Constantly Improving Benchmarks
Jiaxin Ge, Grace Luo, Heekyung Lee, Nishant Malpani, Long Lian, Xudong Wang, Aleksander Holynski, Trevor Darrell, Sewon Min, David M. Chan
摘要
Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these models. Existing benchmarks often lag behind and fail to capture these emerging use cases, leaving a gap between community perceptions of progress and formal evaluation. To address this, we present ECHO, a framework for constructing benchmarks directly from real-world evidence of model use: social media posts that showcase novel prompts and qualitative user judgments. Applying this framework to GPT-4o Image Gen, we construct a dataset of over 31,000 prompts curated from such posts. Our analysis shows that ECHO (1) discovers creative and complex tasks absent from existing benchmarks, such as re-rendering product labels across languages or generating receipts with specified totals, (2) more clearly distinguishes state-of-the-art models from alternatives, and (3) surfaces community feedback that we use to inform the design of metrics for model quality (e.g., measuring observed shifts in color, identity, and structure). Our website is at https://echo-bench.github.io .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
相关 Paper
- DreamBench++: A Human-Aligned Benchmark for Personalized Image GenerationYuang Peng, Yuxin Cui, Haomiao Tang, Zekun Qi 等ICLR 2025
- OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editingzhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu 等ICML 2026 · 被引用 24 次
- PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance PredictionEduard Gabriel Poesina, Adriana Valentina Costache, Adrian-Gabriel Chifu, Josiane Mothe 等CVPR 2025
- Evaluating Image Hallucination in Text-to-Image Generation with Question-AnsweringYoungsun Lim, Hojun Choi, Hyunjung ShimAAAI 2025 · 被引用 16 次
- AutoBencher: Towards Declarative Benchmark ConstructionXiang Lisa Li, Farzaan Kaiyom, Evan Zheran Liu, Yifan Mai 等ICLR 2025
