Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation
Mayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Shin'ichi Satoh
Abstract
Human evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works rely solely on automatic measures (e.g., FID) or perform poorly described human evaluations that are not reliable or repeatable. This paper proposes a standardized and well-defined human evaluation protocol to facilitate verifiable and reproducible human evaluation in future works. In our pilot data collection, we experimentally show that the current automatic measures are incompatible with human perception in evaluating the performance of the text-to-image generation results. Furthermore, we provide insights for designing human evaluation experiments reliably and conclusively. Finally, we make several resources publicly available to the community to facilitate easy and fast implementations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86a0c848-2560-442f-a0cb-b0bee7e7ec80Cited by top-tier papers15
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion modelsGeorge Stein, Jesse C. Cresswell, Rasa Hosseinzadeh, Yi Sui et al.NeurIPS 2023 · 260 citations
- Fair Text-to-Image Diffusion via Fair MappingJia Li, Lijie Hu, Jingfeng Zhang, Tianhang Zheng et al.AAAI 2025 · 36 citations
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text InterpretationJiarui Wang, Huiyu Duan, Ziheng Jia, Zicheng Zhang et al.ICML 2026 · 14 citations
- One-Step Offline Distillation of Diffusion-based Models via Koopman ModelingNimrod Berman, Ilan Naiman, Moshe Eliasof, Hedi Zisling et al.NeurIPS 2025 · 10 citations
Builds on24
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng et al.NeurIPS 2021 · 1,026 citations
Related papers
- ImagenHub: Standardizing the evaluation of conditional image generation modelsMax Ku, Tianle Li, Kai Zhang, Yujie Lu et al.ICLR 2024 · 71 citations
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang et al.ICCV 2023 · 400 citations
- Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human ratingOlivia Wiles, Chuhan Zhang, Isabela Albuquerque, Ivana Kajic et al.ICLR 2025
- Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text GenerationKatelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi, Zongwan Cao et al.ACL 2026
- GENIE: Toward Reproducible and Standardized Human Evaluation for Text GenerationDaniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie et al.EMNLP 2022 · 13 citations
