FAIEr: Fidelity and Adequacy Ensured Image Caption Evaluation
Sijin Wang, Ziwei Yao, Ruiping Wang, Zhongqin Wu, Xilin Chen
摘要
Image caption evaluation is a crucial task, which involves the semantic perception and matching of image and text. Good evaluation metrics aim to be fair, comprehensive, and consistent with human judge intentions. When humans evaluate a caption, they usually consider multiple aspects, such as whether it is related to the target image without distortion, how much image gist it conveys, as well as how fluent and beautiful the language and wording is. The above three different evaluation orientations can be summarized as fidelity, adequacy, and fluency. The former two rely on the image content, while fluency is purely related to linguistics and more subjective. Inspired by human judges, we propose a learning-based metric named FAIEr to ensure evaluating the fidelity and adequacy of the captions. Since image captioning involves two different modalities, we employ the scene graph as a bridge between them to represent both images and captions. FAIEr mainly regards the visual scene graph as the criterion to measure the fidelity. Then for evaluating the adequacy of the candidate caption, it highlights the image gist on the visual scene graph under the guidance of the reference captions. Comprehensive experimental results show that FAIEr has high consistency with human judgment as well as high stability, low reference dependency, and the capability of reference-free evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- Injecting Semantic Concepts into End-to-End Image CaptioningZhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang 等CVPR 2022 · 被引用 125 次
- EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding MatchingYaya Shi, Xu Yang, Haiyang Xu, Chunfeng Yuan 等CVPR 2022 · 被引用 31 次
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- InfoMetIC: An Informative Metric for Reference-free Image Caption EvaluationAnwen Hu, Shizhe Chen, Liang Zhang, Qin JinACL 2023 · 被引用 8 次
它引用的顶会 Paper2
相关 Paper
- LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsShinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki 等AAAI 2026 · 被引用 2 次
- HICEScore: A Hierarchical Metric for Image Captioning EvaluationZequn Zeng, Jianqiao Sun, Hao Zhang, Tiansheng Wen 等ACM MM 2024 · 被引用 3 次
- FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal ModelYebin Lee, Imseong Park, Myungjoo KangACL 2024 · 被引用 3 次
- Cross2StrA: Unpaired Cross-lingual Image Captioning with Cross-lingual Cross-modal Structure-pivoted AlignmentShengqiong Wu, Hao Fei, Wei Ji, Tat-Seng ChuaACL 2023 · 被引用 42 次
- Topic Scene Graph Generation by Attention Distillation from CaptionWenbin Wang, Ruiping Wang, Xilin ChenICCV 2021 · 被引用 16 次
