SMURF: SeMantic and linguistic UndeRstanding Fusion for Caption Evaluation via Typicality Analysis
Joshua Feinglass, Yezhou Yang
Abstract
The open-ended nature of visual captioning makes it a challenging area for evaluation. The majority of proposed models rely on specialized training to improve human-correlation, resulting in limited adoption, generalizability, and explainabilty. We introduce "typicality", a new formulation of evaluation rooted in information theory, which is uniquely suited for problems lacking a definite ground truth. Typicality serves as our framework to develop a novel semantic comparison, SPARCS, as well as referenceless fluency evaluation metrics. Over the course of our analysis, two separate dimensions of fluency naturally emerge: style, captured by metric SPURTS, and grammar, captured in the form of grammatical outlier penalties. Through extensive experiments and ablation studies on benchmark datasets, we show how these decomposed dimensions of semantics and fluency provide greater systemlevel insight into captioner differences. Our proposed metrics along with their combination, SMURF, achieve state-of-the-art correlation with human judgment when compared with other rule-based evaluation metrics 1 . 1 SMURF source codes and data will be released at https: //github.com/JoshuaFeinglass/SMURF .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Injecting Semantic Concepts into End-to-End Image CaptioningZhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang et al.CVPR 2022 · 125 citations
- Polos: Multimodal Metric Learning from Human Feedback for Image CaptioningYuiga Wada, Kanta Kaneda, Daichi Saito, Komei SugiuraCVPR 2024 · 16 citations
- Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation MetricsElisa Kreiss, Cynthia L. Bennett, Shayan Hooshmand, Eric Zelikman et al.EMNLP 2022 · 12 citations
- LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsShinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki et al.AAAI 2026 · 2 citations
Builds on7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Distilling Knowledge Learned in BERT for Text GenerationYen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu et al.ACL 2020 · 116 citations
- Improving Image Captioning Evaluation by Considering Inter References VarianceYanzhi Yi, Hangyu Deng, Jinglu HuACL 2020 · 44 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Perception Score: A Learned Metric for Open-ended Text Generation EvaluationJing Gu, Qingyang Wu, Zhou YuAAAI 2021 · 3 citations
Related papers
- SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption EvaluationXiaofu Chen, Israfel Salazar, Yova KementchedjhievaEMNLP 2025
- FAIEr: Fidelity and Adequacy Ensured Image Caption EvaluationSijin Wang, Ziwei Yao, Ruiping Wang, Zhongqin Wu et al.CVPR 2021
- InfoMetIC: An Informative Metric for Reference-free Image Caption EvaluationAnwen Hu, Shizhe Chen, Liang Zhang, Qin JinACL 2023 · 8 citations
- FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal ModelYebin Lee, Imseong Park, Myungjoo KangACL 2024 · 3 citations
- HICEScore: A Hierarchical Metric for Image Captioning EvaluationZequn Zeng, Jianqiao Sun, Hao Zhang, Tiansheng Wen et al.ACM MM 2024 · 3 citations
