EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding Matching
Yaya Shi, Xu Yang, Haiyang Xu, Chunfeng Yuan, Bing Li, Weiming Hu, Zheng-Jun Zha
Abstract
Current metrics for video captioning are mostly based on the text-level comparison between reference and candidate captions. However, they have some insuperable drawbacks, e.g., they cannot handle videos without references, and they may result in biased evaluation due to the one-to-many nature of video-to-text and the neglect of visual relevance. From the human evaluator's viewpoint, a high-quality caption should be consistent with the provided video, but not necessarily be similar to the reference in literal or semantics. Inspired by human evaluation, we propose EMScore (Embedding Matching-based score), a novel reference-free metric for video captioning, which directly measures similarity between video and candidate captions. Benefiting from the recent development of large-scale pre-training models, we exploit a well pre-trained vision-language model to extract visual and linguistic embeddings for computing EMScore. Specifically, EMScore combines matching scores of both coarse-grained (video and caption) and fine-grained (frames and words) levels, which takes the overall understanding and detailed characteristics of the video into account. Furthermore, considering the potential information gain, EMScore can be flexibly extended to the conditions where human-labeled references are available. Last but not least, we collect VATEX-EVAL and ActivityNet-FOIl datasets to systematically evaluate the existing metrics. VATEX-EVAL experiments demonstrate that EMScore has higher human correlation and lower reference dependency. ActivityNet-FOIL experiment verifies that EMScore can effectively identify “hallucinating” captions. Code and datasets are available at https://github.com/shiyaya/emscore.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4oTony Cheng Tong, Sirui He, Zhiwen Shao, Dit-Yan YeungAAAI 2025 · 22 citations
- Movie101: A New Movie Understanding BenchmarkZihao Yue, Qi Zhang, Anwen Hu, Liang Zhang et al.ACL 2023 · 11 citations
- Better than Random: Reliable NLG Human Evaluation with Constrained Active SamplingJie Ruan, Xiao Pu, Mingqi Gao, Xiaojun Wan et al.AAAI 2024 · 8 citations
- Models See Hallucinations: Evaluating the Factuality in Video CaptioningHui Liu, Xiaojun WanEMNLP 2023 · 5 citations
- Edit As You Wish: Video Caption Editing with Multi-grained User ControlLinli Yao, Yuanmeng Zhang, Ziheng Wang, Xinglin Hou et al.ACM MM 2024 · 4 citations
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- HICEScore: A Hierarchical Metric for Image Captioning EvaluationZequn Zeng, Jianqiao Sun, Hao Zhang, Tiansheng Wen et al.ACM MM 2024 · 3 citations
- VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual AnalysisShubhashis Roy Dipta, Tz-Ying Wu, Subarna TripathiACL 2026
- LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsShinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki et al.AAAI 2026 · 2 citations
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen et al.CVPR 2024 · 38 citations
- Positive-Augmented Contrastive Learning for Image and Video Captioning EvaluationSara Sarto, Manuele Barraco, Marcella Cornia, Lorenzo Baraldi et al.CVPR 2023
