Fine-Grained Multi Image Object Hallucination Benchmark
Joonki Min, Chaeyun Kim, Hyungwook Choi, Yejin Kim, Kihyun Kim, Yohan Jo, Joonseok Lee
摘要
Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination-generating plausible yet factually inconsistent descriptions about objects. Existing benchmarks, designed primarily for single-image settings or providing only high-level multi-image assessments, cannot systematically diagnose how visual complexity and reasoning demands trigger hallucination. To address this gap, we introduce MIOH, a fine-grained multi-image object hallucination benchmark that systematically evaluates object hallucination across four foundational tasks (existence, counting, attribute, position) through three multi-image reasoning patterns (comprehensive, comparative, selective) under three controlled adversarial pressures (visual context scale, perceptual difficulty, contextual bias). Through evaluation of 29 models, we reveal that even state-of-the-art systems like GPT-5 and Gemini-2.5-Pro exhibit distinct failure patterns across different reasoning patterns and tasks. Our evaluation reveals that hallucination stems not merely from perceptual failures but from integration-stage limitations when maintaining object representations across multiple images. MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language ModelsJiale Li, Mingrui Wu, Zixiang Jin, Hao Chen 等ACM MM 2025 · 被引用 4 次
- Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed InputsPeng Ding, Jingyu Wu, Jun Kuang, Dan Ma 等ACM MM 2024 · 被引用 8 次
- Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image SequencesXiyao Wang, Yuhang Zhou, Xiaoyu Liu, Hongjin Lu 等ACL 2024
- PhD: A ChatGPT-Prompted Visual Hallucination Evaluation DatasetJiazhen Liu, Yuhan Fu, Ruobing Xie, Runquan Xie 等CVPR 2025
- THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud ForensicsTzu-Yen Ma, Bo Zhang, Zichen Tang, Junpeng Ding 等ICLR 2026
