MICE-Bench: A Challenging and Comprehensive Benchmark for Multi-Reference Image Creation and Editing
Siqi Luo, Huayu Zheng, Jianghan Shen, Yi Xin, Luxin Xu, Jiyao Liu, Xinyu Zhang, Hang Zhou, Pengyu Xie, Xiaohui Li, Shuo Cao, Yuandong Pu
Abstract
The paradigm of visual generation is rapidly shifting from single-image conditioning toward multi-image conditioning, making the ability to synthesize and edit images based on multiple visual references a critical capability. Despite this trend, existing benchmarks remain largely limited to single-reference scenarios or narrowly defined tasks, leaving model behavior under complex multi-concept composition insufficiently explored. To bridge this gap, we introduce MICE-Bench , a comprehensive benchmark for M ulti-reference I mage C reation and E diting. The benchmark is designed around three core principles: 1) heterogeneous concept composition across seven visual dimensions; 2) varying levels of constraint density, ranging from dual-concept to seven-concept configurations; 3) concept-centric data construction and benchmark evaluation, enabling fine-grained analysis of interactions among multiple concepts. MICE-Bench consists of 3,119 high-quality test cases within a unified concept space. Using an 8-dimensional evaluation metric, we systematically evaluate 13 state-of-the-art models. Our results show that although closed-source models maintain a clear performance advantage, all models experience notable degradation in concept consistency and physical realism as concept complexity increases.This indicates that current models rely on superficial composition rather than genuine multi-concept synthesis, highlighting substantial room for future improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97359c3c-6812-4b06-978d-b1dae353e5bfBuilds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- OmniGen2: Towards Instruction-Aligned Multimodal GenerationChenyuan Wu, Jiahao Wang, Pengfei Zheng, Ruiran Yan et al.CVPR 2026 · 231 citations
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
- I2EBench: A Comprehensive Benchmark for Instruction-based Image EditingYiwei Ma, Jiayi Ji, Ke Ye, Weihuang Lin et al.NeurIPS 2024 · 67 citations
Related papers
- MICo-150K: A Comprehensive Dataset Advancing Multi-Image CompositionXinyu Wei, Kangrui Cen, Hongyang Wei, Zhen Guo et al.CVPR 2026 · 10 citations
- MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image GenerationYuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa et al.CVPR 2026 · 10 citations
- ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and EditingYulin Pan, Xiangteng He, Chaojie Mao, Zhen Han et al.ICCV 2025 · 3 citations
- I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing ModelsJuntong Wang, Jiarui Wang, Huiyu Duan, Jiaxiang Kang et al.CVPR 2026 · 9 citations
- MVGBench: A Comprehensive Benchmark for Multi-View Generation ModelsXianghui Xie, Jan Eric Lenssen, Gerard Pons-MollICCV 2025 · 2 citations
