Lune

EMNLP2024顶会

Holistic Evaluation for Interleaved Text-and-Image Generation

Minqian Liu, Zhiyang Xu, Zihao Lin, Trevor Ashby, Joy Rimchala, Jiaxin Zhang, Lifu Huang

2024年份
1被引次数
16顶会引用

摘要

Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order.Despite the emerging advancements in interleaved generation, the progress in its evaluation still significantly lags behind.Existing evaluation benchmarks do not support arbitrarily interleaved images and text for both inputs and outputs, and they only cover a limited number of domains and use cases.Also, current works predominantly use similarity-based metrics which fall short in assessing the quality in open-ended scenarios.To this end, we introduce INTER-LEAVEDBENCH, the first benchmark carefully curated for the evaluation of interleaved textand-image generation.INTERLEAVEDBENCH features a rich array of tasks to cover diverse real-world use cases.In addition, we present INTERLEAVEDEVAL, a strong reference-free metric powered by GPT-4o to deliver accurate and explainable evaluation.We carefully define five essential evaluation aspects for IN-TERLEAVEDEVAL, including text quality, perceptual quality, image coherence, text-image coherence, and helpfulness, to ensure a comprehensive and fine-grained assessment.Through extensive experiments and rigorous human evaluation, we show that our benchmark and metric can effectively evaluate the existing models with a strong correlation with human judgments surpassing previous reference-based metrics.We also provide substantial findings and insights to foster future research in interleaved generation and its evaluation. 1

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext b1ff11f2-13b1-4e2a-a725-e5898cb7f041

引用它的顶会 Paper16

问问它们各自怎么用它

它引用的顶会 Paper16

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖