Lune

ICCV2025顶会

LMM4LMM: Benchmarking and Evaluating Large-Multimodal Image Generation With LMMs

Jiarui Wang, Huiyu Duan, Yu Zhao, Juntong Wang, Guangtao Zhai, Xiongkuo Min

2025年份
12被引次数
3顶会引用

摘要

2100 (a) Step 1: Prompt collection (b) Step 2: T2I Generation (c) Step 3: Human annotation 16 Annotators 50,400 Generated images <image> <prompt> "a blue cow" Perception (clarity, authenticity, aesthetics) T2I Correspondence (text-image alignment) Does the image contain a cow in the color blue? 0 5 0 5 Yes No 100,800 MOSs from 2 evaluation perspectives 50,400 Yes or no question answering pairs (d) Step 4: Model Design LMM4LMM <A photo of 4 boats> Image 24 LMM-T2I models 50K Images Generated from 2100 Prompts 100K MOSs Assessed from 2 Perspectives 50K Yes or No Question Answer Pairs on 20 Tasks (e) Step 5: Model Comparison LMM EvalMi-50k How would you rate the Perception quality of this image? How would you rate the Correspondence of this image and its prompt? <Prompt> Does the image contain 4 boats? Answer yes or no. A1: The Perception quality of the image is good. Score: 58.21 A2: The Correspondence of the image and its prompt is Bad. Score: 33.33 A3: No. The image does not contain 4 boats. 20 Fine-grained tasks Figure 1. We present the large multimodal image generation evaluation database and model, termed EvalMi-50K and LMM4LMM, respectively. (a) We first collect 2100 comprehensive prompts across 20 fine-grained tasks. (b) Then 24 LMM-T2I models are applied to generate 50K images. (c) 100K MOSs and 50K question-answering pairs are acquired from 16 annotators. (d) We design LMM4LMM to evaluate LMM-T2I models. (e) We conduct model comparisons on EvalMi-50K and the other 7 benchmarks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper32

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖