LMM4LMM: Benchmarking and Evaluating Large-Multimodal Image Generation With LMMs
Jiarui Wang, Huiyu Duan, Yu Zhao, Juntong Wang, Guangtao Zhai, Xiongkuo Min
摘要
2100 (a) Step 1: Prompt collection (b) Step 2: T2I Generation (c) Step 3: Human annotation 16 Annotators 50,400 Generated images <image> <prompt> "a blue cow" Perception (clarity, authenticity, aesthetics) T2I Correspondence (text-image alignment) Does the image contain a cow in the color blue? 0 5 0 5 Yes No 100,800 MOSs from 2 evaluation perspectives 50,400 Yes or no question answering pairs (d) Step 4: Model Design LMM4LMM <A photo of 4 boats> Image 24 LMM-T2I models 50K Images Generated from 2100 Prompts 100K MOSs Assessed from 2 Perspectives 50K Yes or No Question Answer Pairs on 20 Tasks (e) Step 5: Model Comparison LMM EvalMi-50k How would you rate the Perception quality of this image? How would you rate the Correspondence of this image and its prompt? <Prompt> Does the image contain 4 boats? Answer yes or no. A1: The Perception quality of the image is good. Score: 58.21 A2: The Correspondence of the image and its prompt is Bad. Score: 33.33 A3: No. The image does not contain 4 boats. 20 Fine-grained tasks Figure 1. We present the large multimodal image generation evaluation database and model, termed EvalMi-50K and LMM4LMM, respectively. (a) We first collect 2100 comprehensive prompts across 20 fine-grained tasks. (b) Then 24 LMM-T2I models are applied to generate 50K images. (c) 100K MOSs and 50K question-answering pairs are acquired from 16 annotators. (d) We design LMM4LMM to evaluate LMM-T2I models. (e) We conduct model comparisons on EvalMi-50K and the other 7 benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text InterpretationJiarui Wang, Huiyu Duan, Ziheng Jia, Zicheng Zhang 等ICML 2026 · 被引用 14 次
- I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing ModelsJuntong Wang, Jiarui Wang, Huiyu Duan, Jiaxiang Kang 等CVPR 2026 · 被引用 9 次
- VisualScore: Learning Holistic Visual Quality Scores via Multi-Task ReasoningYiting Lu, Fengbin Guan, Yixin Gao, Yan Zhong 等ICML 2026
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Multi-Dimensional Text-to-Face Image Quality Assessment Using LLM: Database and MethodYixuan Gao, Xiongkuo Min, Jinliang Han, Yuqin Cao 等ACM MM 2025 · 被引用 2 次
- What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference AlignmentRishab Parthasarathy, Jasmine Collins, Cory StephensonAAAI 2026
- MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMsYusu Qian, Hanrong Ye, Jean-Philippe Fauconnier, Peter Grasch 等ICLR 2025
- AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics PerceptionYipo Huang, Xiangfei Sheng, Zhichao Yang, Quan Yuan 等ACM MM 2024 · 被引用 34 次
- A High Quality Dataset and Reliable Evaluation for Interleaved Image-Text GenerationYukang Feng, Jianwen Sun, Chuanhao Li, Zizhen Li 等ICLR 2026 · 被引用 4 次
