Holistic Evaluation for Interleaved Text-and-Image Generation
Minqian Liu, Zhiyang Xu, Zihao Lin, Trevor Ashby, Joy Rimchala, Jiaxin Zhang, Lifu Huang
摘要
Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order.Despite the emerging advancements in interleaved generation, the progress in its evaluation still significantly lags behind.Existing evaluation benchmarks do not support arbitrarily interleaved images and text for both inputs and outputs, and they only cover a limited number of domains and use cases.Also, current works predominantly use similarity-based metrics which fall short in assessing the quality in open-ended scenarios.To this end, we introduce INTER-LEAVEDBENCH, the first benchmark carefully curated for the evaluation of interleaved textand-image generation.INTERLEAVEDBENCH features a rich array of tasks to cover diverse real-world use cases.In addition, we present INTERLEAVEDEVAL, a strong reference-free metric powered by GPT-4o to deliver accurate and explainable evaluation.We carefully define five essential evaluation aspects for IN-TERLEAVEDEVAL, including text quality, perceptual quality, image coherence, text-image coherence, and helpfulness, to ensure a comprehensive and fine-grained assessment.Through extensive experiments and rigorous human evaluation, we show that our benchmark and metric can effectively evaluate the existing models with a strong correlation with human judgments surpassing previous reference-based metrics.We also provide substantial findings and insights to foster future research in interleaved generation and its evaluation. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and ImageYushi Hu, Reyhane Askari Hemmat, Melissa Hall, Emily Dinan 等CVPR 2026 · 被引用 18 次
- UniM: A Unified Any-to-Any Interleaved Multimodal BenchmarkYanlin Li, Minghui Guo, Kaiwen Zhang, Shize Zhang 等CVPR 2026 · 被引用 10 次
- Towards Unified Multimodal Interleaved Generation via Group Relative Policy OptimizationMing Nie, Chunwei Wang, Jianhua Han, Hang Xu 等NeurIPS 2025 · 被引用 7 次
- A High Quality Dataset and Reliable Evaluation for Interleaved Image-Text GenerationYukang Feng, Jianwen Sun, Chuanhao Li, Zizhen Li 等ICLR 2026 · 被引用 4 次
- Wan-Weaver: Interleaved Multi-modal Generation via Decoupled TrainingJinbo Xing, Zeyinzi Jiang, Yuxiang Tuo, Chaojie Mao 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper16
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text GenerationPengfei Zhou, Xiaopeng Peng, Jiajun Song, Chuanhao Li 等CVPR 2025
- Interleaved Scene Graphs for Interleaved Text-and-Image Generation AssessmentDongping Chen, Ruoxi Chen, Shu Pu, Zhaoyi Liu 等ICLR 2025
- ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and EditingYulin Pan, Xiangteng He, Chaojie Mao, Zhen Han 等ICCV 2025 · 被引用 3 次
- MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsPeng Xia, Siwei Han, Shi Qiu, Yiyang Zhou 等ICLR 2025
- FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect GeneralizationShibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang 等ICLR 2026 · 被引用 2 次
