Lune

ACL2024顶会

FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model

Yebin Lee, Imseong Park, Myungjoo Kang

2024年份
3被引次数
22顶会引用

摘要

Most existing image captioning evaluation metrics focus on assigning a single numerical score to a caption by comparing it with reference captions.However, these methods do not provide an explanation for the assigned score.Moreover, reference captions are expensive to acquire.In this paper, we propose FLEUR 1 , an explainable reference-free metric to introduce explainability into image captioning evaluation metrics.By leveraging a large multimodal model, FLEUR can evaluate the caption against the image without the need for reference captions, and provide the explanation for the assigned score.We introduce score smoothing to align as closely as possible with human judgment and to be robust to user-defined grading criteria.FLEUR achieves high correlations with human judgment across various image captioning evaluation benchmarks and reaches state-of-the-art results on Flickr8k-CF, COMPOSITE, and Pascal-50S within the domain of reference-free evaluation metrics.Our source code and results are publicly available at: https://github.com/ Yebin46/FLEUR.* Equal contribution.Correspondence to: Myungjoo Kang 1 We choose a word in French that means 'flower', in line with other French-named evaluation metrics.2 A reference caption refers to the human-annotated caption for an image.A candidate caption refers to the caption that is to be evaluated.Score: 0

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper22

问问它们各自怎么用它

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖