EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
Tianyu Chen, Yasi Zhang, Zhi Zhang, Peiyu Yu, Shu Wang, Zhendong Wang, Kevin Lin, Xiaofei Wang, Zhengyuan Yang, Linjie Li, Chung-Ching Lin, Jianwen Xie
Abstract
Instruction-based image editing has advanced rapidly, yet reliable and interpretable evaluation remains a bottleneck. Current protocols either (i) depend on paired reference images—resulting in limited coverage and inheriting biases from prior generative models—or (ii) rely solely on zero-shot vision–language models (VLMs), whose prompt-based assessments of instruction following, content consistency, and visual quality are often imprecise.
To address this, we introduce EdiVal-Agent , an automated and fine-grained evaluation framework grounded in an object-centric perspective, designed to assess not only standard single-turn but also multi-turn instruction-based editing with precision. Given an input image, EdiVal-Agent first decomposes it into semantically meaningful objects, then synthesizes diverse, context-aware editing instructions while dynamically updating object pools across turns. These two stages enable two novel object-centric metrics tailored for multi-turn evaluation and one global metric of visual quality: 1) EdiVal-IF, which measures instruction following by combining open-vocabulary object detectors for symbolic checks with VLMs for semantic verification on detector-guided crops; 2) EdiVal-CC, which evaluates content consistency by calculating semantic similarity of unchanged objects and background using the evolving object pools; and 3) EdiVal-VQ, which quantifies changes in overall visual quality with human preference models.
Instantiating this pipeline, we build EdiVal-Bench, a multi-turn editing benchmark covering 9 instruction types and 16 state-of-the-art editing models spanning in-context, flow-matching, and diffusion paradigms. Our results show that Seedream 4.0 achieves the best overall performance, offering the strongest balance of instruction following, content consistency, and latency. GPT-Image-1.5 clearly improves over GPT-Image-1, especially in content consistency across turns, while Nano Banana 2 consistently outperforms Nano Banana in instruction following and overall score,. Among flow-matching models, FLUX.2-max is the strongest baseline, whereas Qwen-Image-Edit performs well on the first turn but degrades sharply in later turns, indicating strong exposure bias in multi-turn editing. We demonstrate that EdiVal-Agent can be used to identify existing failure modes, thereby informing the development of the next generation of editing models. Our code is available at https://github.com/TianyuCodings/EdiVal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24432c95-cb78-4994-94fe-b380b9e0a8b8Cited by top-tier papers4
- Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object SegmentationHaichao Jiang, Tianming Liang, Wei-Shi Zheng, Jian-Fang HuCVPR 2026 · 7 citations
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image EditingYucheng Liao, Jiajun Liang, Kaiqian Cui, Baoquan Zhao et al.CVPR 2026 · 6 citations
- CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex InstructionsChonghuinan Wang, Zihan Chen, Yuxiang Wei, Tianyi Jiang et al.CVPR 2026 · 3 citations
- Score Distillation Beyond Acceleration: Generative Modeling from Corrupted DataYasi Zhang, Tianyu Chen, Zhendong Wang, Ying Nian Wu et al.ICLR 2026 · 2 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
Related papers
- Image Editing As Programs with Diffusion ModelsYujia Hu, Songhua Liu, Zhenxiong Tan, Xingyi Yang et al.NeurIPS 2025 · 10 citations
- MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial GuidanceXuehai Bai, Xiaoling Gu, Akide Liu, Hangjie Yuan et al.AAAI 2026
- IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing AssessmentYinan Chen, Jiangning Zhang, Teng Hu, Yuxiang Zeng et al.ICLR 2026 · 29 citations
- Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing ModelsYujia Yang, Yuanxiang Wang, Zhenyu Guan, Tiankun Yang et al.CVPR 2026 · 1 citation
- DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing ModelShibo Hong, Boxian Ai, Jun Kuang, Wei Wang et al.ICML 2026
