EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
Ron Yosef, Yonatan Bitton, Dani Lischinski, Moran Yanuka
摘要
Text-guided image editing, fueled by recent advancements in generative AI, is becoming increasingly widespread. This trend highlights the need for a comprehensive framework to verify text-guided edits and assess their quality. To address this need, we introduce EditInspector, a novel benchmark for evaluation of text-guided image edits, based on human annotations collected using an extensive template for edit verification 1 . We leverage EditInspector to evaluate the performance of state-of-the-art (SoTA) vision and language models in assessing edits across various dimensions, including accuracy, artifact detection, visual quality, seamless integration with the image scene, adherence to common sense, and the ability to describe editinduced changes. Our findings indicate that current models struggle to evaluate edits comprehensively and frequently hallucinate when describing the changes. To address these challenges, we propose two novel methods that outperform SoTA models in both artifact detection and difference caption generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Towards Scalable Human-aligned Benchmark for Text-guided Image EditingSuho Ryu, Kihyun Kim, Eugene Baek, Dongsoo Shin 等CVPR 2025
- I2EBench: A Comprehensive Benchmark for Instruction-based Image EditingYiwei Ma, Jiayi Ji, Ke Ye, Weihuang Lin 等NeurIPS 2024 · 被引用 67 次
- Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image InpaintingSu Wang, Chitwan Saharia, Ceslee Montgomery, Jordi Pont-Tuset 等CVPR 2023
- IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing AssessmentYinan Chen, Jiangning Zhang, Teng Hu, Yuxiang Zeng 等ICLR 2026 · 被引用 29 次
- EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing ModelsYupeng Chen, Penglin Chen, Xiaoyu Zhang, Yixian Huang 等AAAI 2025 · 被引用 5 次
