ARGUS: Hallucination and Omission Evaluation in Video-LLMs
Ruchit Rawal, Reza Shirkavand, Heng Huang, Gowthami Somepalli, Tom Goldstein
Abstract
Video large language models have not yet been widely deployed, largely due to their tendency to hallucinate. Typical benchmarks for Video-LLMs rely simply on multiple-choice questions. Unfortunately, VideoLLMs hallucinate far more aggressively on freeform text generation tasks like video captioning than they do on multiple choice verification tasks. To address this weakness, we propose ARGUS, a VideoLLM benchmark that measures freeform video captioning performance. By comparing VideoLLM outputs to human ground truth captions, ARGUS quantifies dual metrics. First, we measure the rate of hallucinations in the form of incorrect statements about video content or temporal relationships. Second, we measure the rate at which the model omits important descriptive details. Together, these dual metrics form a comprehensive view of video captioning performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-FollowingTianyi Xiong, Yi Ge, Ming Li, Zuolong Zhang et al.CVPR 2026 · 16 citations
- Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation EngineeringJianfeng Cai, Jiale Hong, Zongmeng Zhang, Wengang Zhou et al.NeurIPS 2025 · 7 citations
- INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMsJunqi Yang, Yuecong Min, Jie Zhang, Shiguang Shan et al.ACL 2026 · 7 citations
- SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention CollapseYiming Sun, Mi Zhang, Feifei Li, Geng Hong et al.AAAI 2026 · 5 citations
- Catalog-Native LLM: Speaking Item-ID dialect with Less Entanglement for RecommendationReza Shirkavand, Xiaokai Wei, Chen Wang, Zheng Hui et al.ICLR 2026 · 4 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang et al.ICLR 2024 · 476 citations
Related papers
- MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video ModelsGarry Yang, Zizhe Chen, Man Hon Wong, Haoyu Lei et al.ACM MM 2025 · 1 citation
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality EvaluationShi-Xue Zhang, Hongfa Wang, Duojun Huang, Xin Li et al.AAAI 2026 · 5 citations
- Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question AnsweringZizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li et al.ACL 2026
- IF-VidCap: Can Video Caption Models Follow Instructions?Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei et al.ICLR 2026 · 7 citations
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video UnderstandingAshish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand et al.EMNLP 2025 · 1 citation
