INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
Junqi Yang, Yuecong Min, Jie Zhang, Shiguang Shan, Xilin Chen
Abstract
Despite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or verifiable world knowledge (factuality). Existing benchmarks provide limited coverage of factuality hallucinations and predominantly evaluate models only in clean settings. We introduce INFACT, a diagnostic benchmark comprising 9,800 QA instances with fine-grained taxonomies for faithfulness and factuality, spanning real and synthetic videos. INFACT evaluates models in four modes: Base (clean), Visual Degradation, Evidence Corruption, and Temporal Intervention for order-sensitive items. Reliability under induced modes is quantified using Resist Rate (RR) and Temporal Sensitivity Score (TSS). Experiments on 14 representative Video-LLMs reveal that higher Base-mode accuracy does not reliably translate to higher reliability in the induced modes, with evidence corruption reducing stability and temporal intervention yielding the largest degradation. Notably, many open-source baselines exhibit near-zero TSS on factuality, indicating pronounced temporal inertia on order-sensitive questions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d830227-df7c-475e-a78e-68afa87eb74bCited by top-tier papers3
- Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMsWenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang et al.ACL 2026 · 21 citations
- TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image RetrievalZixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen et al.ACL 2026 · 13 citations
- Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented GenerationQianchi Zhang, Hainan Zhang, Liang Pang, Hongwei Zheng et al.ACL 2026 · 3 citations
Builds on9
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- MHBench: Demystifying Motion Hallucination in VideoLLMsMing Kong, Xianzhou Zeng, Luyuan Chen, Yadong Li et al.AAAI 2025 · 6 citations
- ARGUS: Hallucination and Omission Evaluation in Video-LLMsRuchit Rawal, Reza Shirkavand, Heng Huang, Gowthami Somepalli et al.ICCV 2025 · 1 citation
- PhD: A ChatGPT-Prompted Visual Hallucination Evaluation DatasetJiazhen Liu, Yuhan Fu, Ruobing Xie, Runquan Xie et al.CVPR 2025
- MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal RepresentationsKyungho Bae, Jinhyung Kim, Sihaeng Lee, Soonyoung Lee et al.CVPR 2025
Related papers
- MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video ModelsGarry Yang, Zizhe Chen, Man Hon Wong, Haoyu Lei et al.ACM MM 2025 · 1 citation
- Beyond Facts: Evaluating Intent Hallucination in Large Language ModelsYijie Hao, Haofei Yu, Jiaxuan YouACL 2025
- Video SimpleQA: Towards Factuality Evaluation in Large Video Language ModelsMeng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu et al.AAAI 2026 · 1 citation
- VISTA: Verification In Sequential Turn-based AssessmentAshley Lewis, Andrew Perrault, Eric Fosler-Lussier, Michael WhiteACL 2026
- FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke et al.ICLR 2025
