INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
Junqi Yang, Yuecong Min, Jie Zhang, Shiguang Shan, Xilin Chen
摘要
Despite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or verifiable world knowledge (factuality). Existing benchmarks provide limited coverage of factuality hallucinations and predominantly evaluate models only in clean settings. We introduce INFACT, a diagnostic benchmark comprising 9,800 QA instances with fine-grained taxonomies for faithfulness and factuality, spanning real and synthetic videos. INFACT evaluates models in four modes: Base (clean), Visual Degradation, Evidence Corruption, and Temporal Intervention for order-sensitive items. Reliability under induced modes is quantified using Resist Rate (RR) and Temporal Sensitivity Score (TSS). Experiments on 14 representative Video-LLMs reveal that higher Base-mode accuracy does not reliably translate to higher reliability in the induced modes, with evidence corruption reducing stability and temporal intervention yielding the largest degradation. Notably, many open-source baselines exhibit near-zero TSS on factuality, indicating pronounced temporal inertia on order-sensitive questions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMsWenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang 等ACL 2026 · 被引用 21 次
- TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image RetrievalZixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen 等ACL 2026 · 被引用 13 次
- Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented GenerationQianchi Zhang, Hainan Zhang, Liang Pang, Hongwei Zheng 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper9
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- MHBench: Demystifying Motion Hallucination in VideoLLMsMing Kong, Xianzhou Zeng, Luyuan Chen, Yadong Li 等AAAI 2025 · 被引用 6 次
- ARGUS: Hallucination and Omission Evaluation in Video-LLMsRuchit Rawal, Reza Shirkavand, Heng Huang, Gowthami Somepalli 等ICCV 2025 · 被引用 1 次
- PhD: A ChatGPT-Prompted Visual Hallucination Evaluation DatasetJiazhen Liu, Yuhan Fu, Ruobing Xie, Runquan Xie 等CVPR 2025
- MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal RepresentationsKyungho Bae, Jinhyung Kim, Sihaeng Lee, Soonyoung Lee 等CVPR 2025
相关 Paper
- MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video ModelsGarry Yang, Zizhe Chen, Man Hon Wong, Haoyu Lei 等ACM MM 2025 · 被引用 1 次
- Beyond Facts: Evaluating Intent Hallucination in Large Language ModelsYijie Hao, Haofei Yu, Jiaxuan YouACL 2025
- Video SimpleQA: Towards Factuality Evaluation in Large Video Language ModelsMeng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu 等AAAI 2026 · 被引用 1 次
- VISTA: Verification In Sequential Turn-based AssessmentAshley Lewis, Andrew Perrault, Eric Fosler-Lussier, Michael WhiteACL 2026
- FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke 等ICLR 2025
