FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
Yebin Lee, Imseong Park, Myungjoo Kang
Abstract
Most existing image captioning evaluation metrics focus on assigning a single numerical score to a caption by comparing it with reference captions.However, these methods do not provide an explanation for the assigned score.Moreover, reference captions are expensive to acquire.In this paper, we propose FLEUR 1 , an explainable reference-free metric to introduce explainability into image captioning evaluation metrics.By leveraging a large multimodal model, FLEUR can evaluate the caption against the image without the need for reference captions, and provide the explanation for the assigned score.We introduce score smoothing to align as closely as possible with human judgment and to be robust to user-defined grading criteria.FLEUR achieves high correlations with human judgment across various image captioning evaluation benchmarks and reaches state-of-the-art results on Flickr8k-CF, COMPOSITE, and Pascal-50S within the domain of reference-free evaluation metrics.Our source code and results are publicly available at: https://github.com/ Yebin46/FLEUR.* Equal contribution.Correspondence to: Myungjoo Kang 1 We choose a word in French that means 'flower', in line with other French-named evaluation metrics.2 A reference caption refers to the human-annotated caption for an image.A candidate caption refers to the caption that is to be evaluated.Score: 0
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddfaf323-8c68-4ccd-99fd-7bc325b42966Cited by top-tier papers22
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Cycle Consistency as Reward: Learning Image-Text Alignment Without Human PreferencesHyojin Bahng, Caroline Chan, Frédo Durand, Phillip IsolaICCV 2025 · 25 citations
- G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4oTony Cheng Tong, Sirui He, Zhiwen Shao, Dit-Yan YeungAAAI 2025 · 22 citations
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding TasksYadong Niu, TIANZI WANG, Heinrich Dinkel, Xingwei Sun et al.ICML 2026 · 11 citations
- Bridging Human and LLM Judgments: Understanding and Narrowing the GapFelipe Maia Polo, Xinhe Wang, Mikhail Yurochkin, Gongjun Xu et al.NeurIPS 2025 · 7 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsShinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki et al.AAAI 2026 · 2 citations
- HICEScore: A Hierarchical Metric for Image Captioning EvaluationZequn Zeng, Jianqiao Sun, Hao Zhang, Tiansheng Wen et al.ACM MM 2024 · 3 citations
- InfoMetIC: An Informative Metric for Reference-free Image Caption EvaluationAnwen Hu, Shizhe Chen, Liang Zhang, Qin JinACL 2023 · 8 citations
- FAIEr: Fidelity and Adequacy Ensured Image Caption EvaluationSijin Wang, Ziwei Yao, Ruiping Wang, Zhongqin Wu et al.CVPR 2021
- VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image CaptionsKazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki et al.EMNLP 2025
