FINER: MLLMs Hallucinate under Fine-grained Negative Queries
Rui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata, Stephan Alaniz
摘要
Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine-grained queries, a challenge underrepresented by existing benchmarks that focus on coarse image-related questions. We introduce FInegrained NEgative queRies (FINER), alongside two benchmarks: FINER-CompreCap and FINER-DOCCI. Using FINER, we analyze hallucinations across four settings: multi-object, multi-attribute, multi-relation, and "what" questions. Our benchmarks reveal that MLLMs hallucinate when fine-grained mismatches co-occur with genuinely present elements in the image. To address this, we propose FINER-Tuning, leveraging Direct Preference Optimization (DPO) on FINER-inspired data. Finetuning four frontier MLLMs with FINER-Tuning yields up to 24.2% gains (InternVL3.5-14B) on hallucinations from our benchmarks, while simultaneously improving performance on eight existing hallucination suites, and enhancing general multimodal capabilities across six benchmarks. Code, benchmark, and models are available at https://explainableml.github.io/finer-project/. Can you see the cat in this image? Can you see the wolf in this image? Can you see the cat with predominantly white coat featuring black and grey markings in this image? Can you see the cat with predominantly brown coat featuring orange and pink markings in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head turned downwards in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head tilted backward in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head turned downwards, with perked ears in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head turned downwards, with drooping ears in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head turned downwards, with perked ears , that is sitting on the chair in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head turned downwards, with perked ears , that is sitting below the chair in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head turned downwards, with perked ears , that is sitting on the chair in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head turned downwards, with perked ears , that is sitting on the sofa in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head turned downwards, with perked ears , that is sitting on the chair with primarily blue and white in color exhibiting signs of wear in this image? Can you see the cat with predominantly white coat featuring black and grey markings, with its head turned downwards, with perked ears , that is sitting on the chair with largely purple and orange in color indicating heavy use in this image? Can you see the cat in this image? Can you see the wolf? Can you see the ca and grey markings i Can you see the cat with predominantly brown coat? Can you see the cat w image? Can you see the cat with …, with its head tilted backward? Can you see the cat with …, …, with perked ears in th Can you see the cat with …, …, with drooping ears? Can you see the cat with …, …, …, that is sitting on the this image? Can you see the cat with …, ..., …, that is sitting below the chair? Can you see the cat with …, …, …, that is sitting on the this image?
Can you see the cat with …, …, …, that is sitting on the sofa?
Can you see the cat with …, …, …, that is … with prima and white in color in this image?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
相关 Paper
- CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMsJinlan Fu, Shenzhen Huangfu, Hao Fei, Xiaoyu Shen 等ICLR 2025
- Adaptive Hallucination Alleviation in Multimodal Large Language Models: From Strategic Data Selection to Severity-Guided TrainingYuanyi Xu, Xiangru Zhu, Sihang Jiang, Zhixu Li 等AAAI 2026
- Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language ModelsJeonghwan Kim, Heng JiEMNLP 2024 · 被引用 4 次
- Cat-PO: Cross-modal Adaptive Token-rewards for Preference Optimization in Truthful Multimodal LLMsZhixiao Zheng, Zheren Fu, Zhiyuan Yao, Dongming Zhang 等ICLR 2026
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image GenerationYoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi 等CVPR 2026
