Models See Hallucinations: Evaluating the Factuality in Video Captioning
Hui Liu, Xiaojun Wan
Abstract
Video captioning aims to describe events in a video with natural language. In recent years, many works have focused on improving captioning models' performance. However, like other text generation tasks, it risks introducing factual errors not supported by the input video. Factual errors can seriously affect the quality of the generated text, sometimes making it completely unusable. Although factual consistency has received much research attention in text-to-text tasks (e.g., summarization), it is less studied in vision-based text generation. In this work, we conduct the first human evaluation of the factuality in video captioning and annotate two factuality datasets. We find that 56% of the model-generated sentences have factual errors, indicating it is a severe problem in this field, but existing evaluation metrics show little correlation with human factuality annotation. We further propose a weakly-supervised, model-based factuality metric FactVC, which outperforms previous metrics on factuality evaluation of video captioning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c2c7a7c-905e-4ea3-8ae5-71a4b680059fCited by top-tier papers5
- An Audit on the Perspectives and Challenges of Hallucinations in NLPPranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs et al.EMNLP 2024 · 8 citations
- Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language TranslationYasser Hamidullah, Koel Dutta Chowdhury, Yusser Al Ghussin, Shakib Yazdani et al.ICLR 2026 · 3 citations
- Summarizing Speech: A Comprehensive SurveyFabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau et al.EMNLP 2025 · 3 citations
- What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific PresentationsDongqi Liu, Chenxi Whitehouse, Xi Yu, Louis Mahon et al.ACL 2025
- VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual AnalysisShubhashis Roy Dipta, Tz-Ying Wu, Subarna TripathiACL 2026
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
Related papers
- Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question AnsweringZizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li et al.ACL 2026
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality EvaluationShi-Xue Zhang, Hongfa Wang, Duojun Huang, Xin Li et al.AAAI 2026 · 5 citations
- EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding MatchingYaya Shi, Xu Yang, Haiyang Xu, Chunfeng Yuan et al.CVPR 2022 · 31 citations
- WeCheck: Strong Factual Consistency Checker via Weakly Supervised LearningWenhao Wu, Wei Li, Xinyan Xiao, Jiachen Liu et al.ACL 2023 · 3 citations
- ARGUS: Hallucination and Omission Evaluation in Video-LLMsRuchit Rawal, Reza Shirkavand, Heng Huang, Gowthami Somepalli et al.ICCV 2025 · 1 citation
