FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
Zhihan Yin, Jianxin Liang, Yueqian Wang, Yifeng Yao, Huishuai Zhang, Dongyan Zhao
Abstract
Multimodal Large Language Models (MLLMs) suffer from hallucinations. Existing hallucination evaluation benchmarks are often limited by over-simplified tasks leading to saturated metrics, or insufficient diversity that fails to adequately assess the hallucination extent in state-of-the-art multimodal models. To address this gap, we propose FREAK, a comprehensive multimodal benchmark designed for fine-grained hallucination assessment in MLLMs. Through high-quality photorealistic images featuring fine-grained counter-commonsense edits, FREAK innovatively evaluates hallucination phenomena in detailed visual perception of MLLMs. Extensive experiments on FREAK show severe hallucination issues in SOTA models regarding detailed visual perception. To enable deeper investigation, we curate a controlled subset to indirectly evaluate the model’s ability to perceive target detailed information. Through systematic evaluation of prevailing Chain-of-Thought (CoT) prompting techniques within this task, we reveal critical insights regarding hallucination patterns and model reasoning processes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3abed3db-c9ff-4f31-a0f5-83142bacedb6Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLMBowen Dong, Minheng Ni, Zitong Huang, Guanglei Yang et al.NeurIPS 2025 · 25 citations
- Fine-Grained Multi Image Object Hallucination BenchmarkJoonki Min, Chaeyun Kim, Hyungwook Choi, Yejin Kim et al.CVPR 2026 · 1 citation
- Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language ModelsChaoya Jiang, Hongrui Jia, Mengfan Dong, Wei Ye et al.ACM MM 2024 · 19 citations
- Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and ReasoningJunpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji et al.ACL 2026
- PhD: A ChatGPT-Prompted Visual Hallucination Evaluation DatasetJiazhen Liu, Yuhan Fu, Ruobing Xie, Runquan Xie et al.CVPR 2025
