T2I-Scorer: Quantitative Evaluation on Text-to-Image Generation via Fine-Tuned Large Multi-Modal Models
Haoning Wu, Xiele Wu, Chunyi Li, Zicheng Zhang, Chaofeng Chen, Xiaohong Liu, Guangtao Zhai, Weisi Lin
摘要
Text-to-image (T2I) generation is a pivotal and core interest within the realm of AI content generation. Amid the swift advancements of both open-source (such as Stable Diffusion) and proprietary (for example, DALLE, MidJourney) T2I models, there is a notable absence of a comprehensive and robust quantitative framework for evaluating their output quality. Traditional methods of quality assessment overlook the textual prompts when judging images; meanwhile, the advent of large multi-modal models (LMMs) introduces the capability to incorporate text prompts in evaluations, yet the challenge of fine-tuning these models for precise T2I quality assessment remains unresolved. In our study, we introduce the T2I-Scorer, a novel two-stage training methodology aimed at fine-tuning LMMs for T2I evaluation. For the first stage, we collect 397K GPT-4V-labeled question-answer pairs related to T2I evaluation. Termed as T2I-ITD, the pseudo-labeled dataset is analyzed and examined by human, and used for instruction tuning to improve the LMM's low-level quality perception. The first stage model, T2I-Scorer-IT, has reached superior accuracy on T2I evaluation than all kinds of existing T2I metrics under zero-shot settings. For the second stage, we define an explicit multi-task training scheme to further align the LMM with human opinion scores, and the fine-tuned T2I-Scorer can reach state-of-the-art accuracy on both image quality and image-text alignment perspectives with significant improvements. We anticipate the proposed metrics can serve as a reliable metric to gauge the ability of T2I generation models in the future. We will make code, data, and weights publicly available.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Cap: Evaluation of Persuasive and Creative Image GenerationAysan Aghazadeh, Adriana KovashkaICCV 2025 · 被引用 9 次
- Adaptive Routing of Text-to-Image Generation Requests between Large Cloud Model and Light-Weight Edge ModelZewei Xin, Qinya Li, Chaoyue Niu, Fan Wu 等ICCV 2025 · 被引用 2 次
- Evaluating Visual Narrative Coherence in Story Visualization via Diversified StorylinesMinha Jhang, Kyeongman Park, Hyukhun Koh, Kyomin JungACL 2026
- VisualScore: Learning Holistic Visual Quality Scores via Multi-Task ReasoningYiting Lu, Fengbin Guan, Yixin Gao, Yan Zhong 等ICML 2026
相关 Paper
- Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation BenchmarkRong-Cheng Tu, Zi-Ao Ma, Tian Lan, Yuehao Zhao 等ACL 2025 · 被引用 13 次
- Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language ModelsChutian Meng, Fan Ma, Jiaxu Miao, Chi Zhang 等AAAI 2025 · 被引用 1 次
- Revisiting MLLM Based Image Quality Assessment: Errors and RemedyZhenchen Tang, Songlin Yang, Bo Peng, Zichuan Wang 等AAAI 2026 · 被引用 2 次
- ScImage: How good are multimodal large language models at scientific text-to-image generation?Leixin Zhang, Steffen Eger, Yinjie Cheng, Weihe Zhai 等ICLR 2025
- Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension ScoringQizhi Xie, Kun Yuan, Yunpeng Qu, Jiachao Gong 等CVPR 2026
