Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
Rong-Cheng Tu, Zi-Ao Ma, Tian Lan, Yuehao Zhao, Heyan Huang, Xian-Ling Mao
摘要
Driven by the remarkable progress in diffusion models, text-to-image generation has made significant strides, creating a pressing demand for automatic quality evaluation of generated images. Current state-of-the-art automatic evaluation methods heavily rely on Multi-modal Large Language Models (MLLMs), particularly powerful commercial models like GPT-4o. While these models are highly effective, their substantial costs limit scalability in large-scale evaluations. Adopting open-source MLLMs is an alternative; however, their performance falls short due to significant limitations in processing multi-modal data compared to commercial MLLMs. To tackle these problems, we first propose a task decomposition evaluation framework based on GPT-4o to automatically construct a new training dataset, where the complex evaluation task is decoupled into simpler sub-tasks, effectively reducing the learning complexity. Based on this dataset, we design innovative training strategies to effectively distill GPT-4o's evaluation capabilities into a 7B open-source MLLM, MiniCPM-V-2.6. Furthermore, to reliably and comprehensively assess prior works and our proposed model, we manually annotate a meta-evaluation benchmark that includes chain-of-thought explanations alongside quality scores for generated images. Experimental results demonstrate that our distilled open-source MLLM significantly outperforms the current state-of-the-art GPT-4o-base baseline, VIEScore, with over 4.6% improvement in Spearman and Kendall correlations with human judgments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Multimodal LLMs as Customized Reward Models for Text-to-Image GenerationShijie Zhou, Ruiyi Zhang, Huaisheng Zhu, Branislav Kveton 等ICCV 2025 · 被引用 3 次
- REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment EvaluationFulin Shi, Wenyi Xiao, Bin Chen, Liang Ding 等ACL 2026 · 被引用 3 次
- Self-Supervised Direct Preference Optimization for Text-to-Image Diffusion ModelsLiang Peng, Boxi Wu, Haoran Cheng, Yibo Zhao 等NeurIPS 2025 · 被引用 2 次
- Engaging Communities Meaningfully in Defining Disability Representation for AI Image GenerationAnja Thieme, Rita Faia Marques, Martin Grayson, Sidhika Balachandar 等CHI 2026 · 被引用 1 次
- Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head GenerationJianzhi Long, Wenhao Sun, Rong-Cheng Tu, Dacheng TaoAAAI 2026
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- VIEScore: Towards Explainable Metrics for Conditional Image Synthesis EvaluationMax Ku, Dongfu Jiang, Cong Wei, Xiang Yue 等ACL 2024 · 被引用 24 次
- T2I-Scorer: Quantitative Evaluation on Text-to-Image Generation via Fine-Tuned Large Multi-Modal ModelsHaoning Wu, Xiele Wu, Chunyi Li, Zicheng Zhang 等ACM MM 2024 · 被引用 8 次
- Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment LearningQinghao Ye, Xianhan Zeng, Fu Li, Chunyuan Li 等ICLR 2025
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable ApproachDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma 等ICLR 2026 · 被引用 4 次
- CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex InstructionsChonghuinan Wang, Zihan Chen, Yuxiang Wei, Tianyi Jiang 等CVPR 2026 · 被引用 3 次
