VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
Chaoyu Li, Eun Woo Im, Pooyan Fazli
摘要
Multimodal large language models (MLLMs) have recently shown significant advancements in video understanding, excelling in content reasoning and instruction-following tasks. However, hallucination, where models generate inaccurate or misleading content, remains underexplored in the video domain. Building on the observation that MLLM visual encoders often fail to distinguish visually different yet semantically similar video pairs, we introduce VIDHAL-LUC, the largest benchmark designed to examine hallucinations in MLLMs for video understanding. It consists of 5,002 videos, paired to highlight cases prone to hallucinations. VIDHALLUC assesses hallucinations across three critical dimensions: (1) action, (2) temporal sequence, and (3) scene transition. Comprehensive testing shows that most MLLMs are vulnerable to hallucinations across these dimensions. Furthermore, we propose DINO-HEAL, a trainingfree method that reduces hallucinations by incorporating spatial saliency from DINOv2 to reweight visual features during inference. Our results show that DINO-HEAL consistently improves performance on VIDHALLUC, achieving an average improvement of 3.02% in mitigating hallucinations across all tasks. Both the VIDHALLUC benchmark and DINO-HEAL code are available at https://peoplerobots.github.io/vidhalluc . Q1) Is the prominent action in the video playing the piano? Q2) Is the prominent action in the video playing drums? Chat-UniVi: Yes, the prominent action in the video is playing the piano.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware ReasoningSARA GHAZANFARI, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy 等CVPR 2026 · 被引用 35 次
- VideoA11y: Method and Dataset for Accessible Video DescriptionChaoyu Li, Sid Padmanabhuni, Maryam S. Cheema, Hasti Seifi 等CHI 2025 · 被引用 23 次
- Time Blindness: Why Video-Language Models Can't See What Humans Can?Ujjwal Upadhyay, Mukul Ranjan, Zhiqiang Shen, Mohamed ElhoseinyCVPR 2026 · 被引用 17 次
- AVATAR: Reinforcement Learning to See, Hear, and Reason Over VideoYogesh Kulkarni, Pooyan FazliCVPR 2026 · 被引用 15 次
- SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive DecodingChang-Hsun Wu, Kai-Po Chang, Yu-Yang Sheng, Hung-Kai Chung 等CVPR 2026 · 被引用 8 次
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video UnderstandingHao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang 等CVPR 2026
- Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language ModelsWenbin Xing, Quanxing Zha, Lizheng Zu, Mengran Li 等ICML 2026 · 被引用 1 次
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video UnderstandingAshish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand 等EMNLP 2025 · 被引用 1 次
- MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video ModelsGarry Yang, Zizhe Chen, Man Hon Wong, Haoyu Lei 等ACM MM 2025 · 被引用 1 次
- VisDiaHalBench: A Visual Dialogue Benchmark For Diagnosing Hallucination in Large Vision-Language ModelsQingxing Cao, Junhao Cheng, Xiaodan Liang, Liang LinACL 2024 · 被引用 3 次
