Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
Meng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu, Haoran Tang, Haoze Zhao, Chen Wang, Jiahua Dong, Wangbo Yu, Ge Zhang, Xiang Li, Ian Reid, Xiaodan Liang
摘要
Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains a critical unsolved challenge. To address this gap, we introduce Video SimpleQA, the first comprehensive benchmark tailored for factuality evaluation in video contexts. Our work differs from existing video benchmarks through the following key features: 1) Knowledge required: demanding integration of external knowledge beyond the video’s explicit narrative; 2) Multi-hop fact-seeking question: Each question involves multiple explicit facts and requires strict factual grounding without hypothetical or subjective inferences. We include per-hop single-fact-based sub-QAs alongside final QAs to enable fine-grained, step-by-step evaluation; 3) Short-form definitive answer: Answers are crafted as unambiguous and definitively correct in a short format with minimal scoring variance; 4) Temporal grounded required: Requiring answers to rely on one or more temporal segments in videos, rather than single frames. We extensively evaluate 33 state-of-the-art LVLMs and summarize key findings as follows: 1) Current LVLMs exhibit notable deficiencies in factual adherence, with the best-performing model o3 merely achieving an F-score of 66.3%; 2) Most LVLMs are overconfident in what they generate, with self-stated confidence exceeding actual accuracy; 3) Retrieval-Augmented Generation demonstrates consistent improvements at the cost of additional inference time overhead; 4) Multi-hop QA demonstrates substantially degraded performance compared to single-hop sub-QAs, with first-hop object/event recognition emerging as the primary bottleneck. We position Video SimpleQA as the cornerstone benchmark for video factuality assessment, aiming to steer LVLM development toward verifiable grounding in real-world contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMsWenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang 等ACL 2026 · 被引用 21 次
- MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering BenchmarkShaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie 等CVPR 2026 · 被引用 4 次
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric ScenesKeliang Li, Hongze Shen, Hao Shi, Ruibing Hou 等ICLR 2026 · 被引用 2 次
- From Detection to Understanding: Multi-Turn Reasoning for Video Misinformation AnalysisZhi Zeng, Jiaying Wu, Minnan Luo, Di Zhang 等ACL 2026
它引用的顶会 Paper16
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen 等ICLR 2024 · 被引用 699 次
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang 等CVPR 2024 · 被引用 95 次
相关 Paper
- SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language ModelsXianfu Cheng, Wei Zhang, Shiwei Zhang, Jian Yang 等ICCV 2025 · 被引用 4 次
- INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMsJunqi Yang, Yuecong Min, Jie Zhang, Shiguang Shan 等ACL 2026 · 被引用 7 次
- Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language ModelsYancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan 等ACL 2025
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang 等NeurIPS 2025 · 被引用 30 次
- Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosQirui Chen, Shangzhe Di, Weidi XieAAAI 2025 · 被引用 35 次
