Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
Tingyu Song, Yanzhao Zhang, Mingxin Li, Zhuoning Guo, Dingkun Long, Pengjun Xie, Siyue Zhang, Yilun Zhao, Shu Wu
摘要
Composed Image Retrieval (CIR) is a pivotal and complex task in multimodal understanding. Current CIR benchmarks typically feature limited query categories and fail to capture the diverse requirements of real-world scenarios. To bridge this evaluation gap, we leverage image editing to achieve precise control over modification types and content, enabling a pipeline for synthesizing queries across a broad spectrum of categories. Using this pipeline, we construct EDIR, a novel fine-grained CIR benchmark. EDIR encompasses 5,000 high-quality queries structured across five main categories and fifteen subcategories. Our comprehensive evaluation of 13 multimodal embedding models reveals a significant capability gap; even state-of-the-art models (e.g., RzenEmbed and GME) struggle to perform consistently across all subcategories, highlighting the rigorous nature of our benchmark. Through comparative analysis, we further uncover inherent limitations in existing benchmarks, such as modality biases and insufficient categorical coverage. Furthermore, an in-domain training experiment demonstrates the feasibility of our benchmark. This experiment clarifies the task challenges by distinguishing between categories that are solvable with targeted data and those that expose the intrinsic limitations of current model architectures 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 被引用 344 次
- Zero-Shot Composed Image Retrieval with Textual InversionAlberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del BimboICCV 2023 · 被引用 214 次
- MagicLens: Self-Supervised Image Retrieval with Open-Ended InstructionsKai Zhang, Yi Luan, Hexiang Hu, Kenton Lee 等ICML 2024 · 被引用 112 次
- Learning Attribute-driven Disentangled Representations for Interactive Fashion RetrievalYuxin Hou, Eleonora Vig, Michael Donoser, Loris BazzaniICCV 2021 · 被引用 58 次
相关 Paper
- MA-CIR: A Multimodal Arithmetic Benchmark for Composed Image RetrievalJaeseok Byun, Young Kyun Jang, Seokhyeon Jeong, Donghyun Kim 等ICCV 2025 · 被引用 1 次
- CoLLM: A Large Language Model for Composed Image RetrievalChuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah 等CVPR 2025
- ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalZixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu 等AAAI 2025 · 被引用 59 次
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou 等CVPR 2025
- RMIR: A Benchmark Dataset for Reasoning-Intensive Multimodal Image RetrievalYijiang Li, Kunal Kotian, Ali Marjaninejad, Meir Friedenberg 等CVPR 2026
