LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
Zitong Xu, Huiyu Duan, Bingnan Liu, Guangji Ma, Jiarui Wang, Liu Yang, Shiqi Gao, Xiaoyu Wang, Jia Wang, Xiongkuo Min, Guangtao Zhai, Weisi Lin
Abstract
The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image, limiting their practical applications. Existing TIE evaluation benchmarks and metrics have limitations on scale or alignment with human perception. To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. Specifically, EBench-18K includes 1,080 source images with corresponding editing prompts across 21 tasks, 18K+ edited images produced by 17 state-of-the-art TIE models, 55K+ mean opinion scores (MOSs) assessed from three evaluation dimensions, and 18K+ question-answering (QA) pairs. Based on EBench-18K, we employ outstanding LMMs to assess edited images, while the evaluation results, in turn, provide insights into assessing the alignment between the LMMs' understanding ability and human preferences. Then, we propose LMM4Edit, a LMM-based metric for evaluating image Editing models from perceptual quality, editing alignment, attribute preservation, and task-specific QA accuracy in an all-in-one manner. Extensive experiments show that LMM4Edit achieves outstanding performance and aligns well with human preference. Zero-shot validation on the other datasets also shows the generalization ability of our model. The dataset and code are available at https://github.com/IntMeGroup/LMM4Edit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image EditingWei Chow, Linfeng Li, Lingdong Kong, Zefeng Li et al.CVPR 2026 · 14 citations
- I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing ModelsJuntong Wang, Jiarui Wang, Huiyu Duan, Jiaxiang Kang et al.CVPR 2026 · 9 citations
- AutoEdit: Automatic Hyperparameter Tuning for Image EditingChau Pham, Quan Dao, Mahesh Bhosale, Yunjie Tian et al.NeurIPS 2025 · 3 citations
- Omni2: Unifying Omnidirectional Image Generation and Editing in an Omni ModelLiu Yang, Huiyu Duan, Yucheng Zhu, Xiaohong Liu et al.ACM MM 2025 · 2 citations
- Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting FrameworkJiayi Gao, Qingchao Chen, Yuxin Peng, Yang LiuICML 2026 · 1 citation
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Towards Scalable Human-aligned Benchmark for Text-guided Image EditingSuho Ryu, Kihyun Kim, Eugene Baek, Dongsoo Shin et al.CVPR 2025
- I2EBench: A Comprehensive Benchmark for Instruction-based Image EditingYiwei Ma, Jiayi Ji, Ke Ye, Weihuang Lin et al.NeurIPS 2024 · 67 citations
- VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality AssessmentShangkun Sun, Xiaoyu Liang, Songlin Fan, Wenxu Gao et al.AAAI 2025 · 19 citations
- Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing ModelsYujia Yang, Yuanxiang Wang, Zhenyu Guan, Tiankun Yang et al.CVPR 2026 · 1 citation
- CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex InstructionsChonghuinan Wang, Zihan Chen, Yuxiang Wei, Tianyi Jiang et al.CVPR 2026 · 3 citations
