Improving Image Captioning Evaluation by Considering Inter References Variance
Yanzhi Yi, Hangyu Deng, Jinglu Hu
Abstract
Evaluating image captions is very challenging partially due to the fact that there are multiple correct captions for every single image. Most of the existing one-to-one metrics operate by penalizing mismatches between reference and generative caption without considering the intrinsic variance between ground truth captions. It usually leads to over-penalization and thus a bad correlation to human judgment. Recently, the latest one-to-one metric BERTScore can achieve high human correlation in system-level tasks while some issues can be fixed for better performance. In this paper, we propose a novel metric based on BERTScore that could handle such a challenge and extend BERTScore with a few new features appropriately for image captioning evaluation. The experimental results show that our metric achieves state-of-the-art human judgment correlation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 746e6d6f-2e0d-491b-a32e-cb2942e3ab3aCited by top-tier papers16
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Mutual Information Divergence: A Unified Metric for Multimodal Generative ModelsJin-Hwa Kim, Yunji Kim, Jiyoung Lee, Kang Min Yoo et al.NeurIPS 2022 · 49 citations
- EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding MatchingYaya Shi, Xu Yang, Haiyang Xu, Chunfeng Yuan et al.CVPR 2022 · 31 citations
- Multi-modal Dependency Tree for Video CaptioningWentian Zhao, Xinxiao Wu, Jiebo LuoNeurIPS 2021 · 27 citations
- Test-Time Distribution Normalization for Contrastively Learned Visual-language ModelsYifei Zhou, Juntao Ren, Fengyu Li, Ramin Zabih et al.NeurIPS 2023 · 22 citations
Builds on3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
Related papers
- InfoMetIC: An Informative Metric for Reference-free Image Caption EvaluationAnwen Hu, Shizhe Chen, Liang Zhang, Qin JinACL 2023 · 8 citations
- Revisiting Grammatical Error Correction Evaluation and BeyondPeiyuan Gong, Xuebo Liu, Heyan Huang, Min ZhangEMNLP 2022 · 11 citations
- HICEScore: A Hierarchical Metric for Image Captioning EvaluationZequn Zeng, Jianqiao Sun, Hao Zhang, Tiansheng Wen et al.ACM MM 2024 · 3 citations
- QRelScore: Better Evaluating Generated Questions with Deeper Understanding of Context-aware RelevanceXiaoqiang Wang, Bang Liu, Siliang Tang, Lingfei WuEMNLP 2022 · 6 citations
- LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsShinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki et al.AAAI 2026 · 2 citations
