LMM4LMM: Benchmarking and Evaluating Large-Multimodal Image Generation With LMMs
Jiarui Wang, Huiyu Duan, Yu Zhao, Juntong Wang, Guangtao Zhai, Xiongkuo Min
Abstract
2100 (a) Step 1: Prompt collection (b) Step 2: T2I Generation (c) Step 3: Human annotation 16 Annotators 50,400 Generated images <image> <prompt> "a blue cow" Perception (clarity, authenticity, aesthetics) T2I Correspondence (text-image alignment) Does the image contain a cow in the color blue? 0 5 0 5 Yes No 100,800 MOSs from 2 evaluation perspectives 50,400 Yes or no question answering pairs (d) Step 4: Model Design LMM4LMM <A photo of 4 boats> Image 24 LMM-T2I models 50K Images Generated from 2100 Prompts 100K MOSs Assessed from 2 Perspectives 50K Yes or No Question Answer Pairs on 20 Tasks (e) Step 5: Model Comparison LMM EvalMi-50k How would you rate the Perception quality of this image? How would you rate the Correspondence of this image and its prompt? <Prompt> Does the image contain 4 boats? Answer yes or no. A1: The Perception quality of the image is good. Score: 58.21 A2: The Correspondence of the image and its prompt is Bad. Score: 33.33 A3: No. The image does not contain 4 boats. 20 Fine-grained tasks Figure 1. We present the large multimodal image generation evaluation database and model, termed EvalMi-50K and LMM4LMM, respectively. (a) We first collect 2100 comprehensive prompts across 20 fine-grained tasks. (b) Then 24 LMM-T2I models are applied to generate 50K images. (c) 100K MOSs and 50K question-answering pairs are acquired from 16 annotators. (d) We design LMM4LMM to evaluate LMM-T2I models. (e) We conduct model comparisons on EvalMi-50K and the other 7 benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4e747d2-f19e-4ef3-b020-879cadae4b0cCited by top-tier papers3
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text InterpretationJiarui Wang, Huiyu Duan, Ziheng Jia, Zicheng Zhang et al.ICML 2026 · 14 citations
- I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing ModelsJuntong Wang, Jiarui Wang, Huiyu Duan, Jiaxiang Kang et al.CVPR 2026 · 9 citations
- VisualScore: Learning Holistic Visual Quality Scores via Multi-Task ReasoningYiting Lu, Fengbin Guan, Yixin Gao, Yan Zhong et al.ICML 2026
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Multi-Dimensional Text-to-Face Image Quality Assessment Using LLM: Database and MethodYixuan Gao, Xiongkuo Min, Jinliang Han, Yuqin Cao et al.ACM MM 2025 · 2 citations
- What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference AlignmentRishab Parthasarathy, Jasmine Collins, Cory StephensonAAAI 2026
- MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMsYusu Qian, Hanrong Ye, Jean-Philippe Fauconnier, Peter Grasch et al.ICLR 2025
- AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics PerceptionYipo Huang, Xiangfei Sheng, Zhichao Yang, Quan Yuan et al.ACM MM 2024 · 34 citations
- A High Quality Dataset and Reliable Evaluation for Interleaved Image-Text GenerationYukang Feng, Jianwen Sun, Chuanhao Li, Zizhen Li et al.ICLR 2026 · 4 citations
