Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
Pritam Sarkar, Sayna Ebrahimi, Ali Etemad, Ahmad Beirami, Sercan Ö. Arik, Tomas Pfister
摘要
Despite their significant advancements, Multimodal Large Language Models (MLLMs) often generate factually inaccurate information, referred to as hallucination. In this work, we address object hallucinations in MLLMs, where information is generated about an object not present in the input image. We introduce Dataaugmented Phrase-level Alignment (DPA), a novel loss which can be applied to instruction-tuned off-the-shelf MLLMs to mitigate hallucinations, while preserving their general vision-language capabilities. To fine-tune MLLMs with DPA, we first generate a set of 'hallucinated' and 'correct' response pairs through generative data augmentation by selectively altering the ground-truth information of the correct responses at a phrase level. The DPA loss is then used to train MLLMs to reduce the likelihood of hallucinated phrases compared to the correct ones. In contrast, existing alignment techniques act at the sequence level and often lead to a sharp trade off between mitigating hallucinations and preserving model capabilities. Our thorough evaluation on various benchmarks confirms the effectiveness of DPA in mitigating hallucination while retaining the out-of-the-box performance of the MLLMs on general tasks. For instance, MLLMs finetuned with DPA, which we refer to as Hallucination Attenuated Language and Vision Assistant (HALVA), improve F1 by up to 13.4% on hallucination visual question-answering and reduce the hallucination rate by up to 4.2% on image description tasks. LLaVA-v1.5 HALVA In the image, there are several utensils, including forks, knives, and spoons, made out of Legos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLMBowen Dong, Minheng Ni, Zitong Huang, Guanglei Yang 等NeurIPS 2025 · 被引用 25 次
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference OptimizationWenqi Liu, Xuemeng Song, Jiaxi Li, Yinwei Wei 等NeurIPS 2025 · 被引用 18 次
- Co-Reinforcement Learning for Unified Multimodal Understanding and GenerationJingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang 等NeurIPS 2025 · 被引用 15 次
- MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language ModelsQiyan Zhao, Xiaofeng Zhang, Yiheng Li, Yun Xing 等ACM MM 2025 · 被引用 11 次
- SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual ScenesChuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu 等ACL 2026 · 被引用 9 次
它引用的顶会 Paper36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionWenbin An, Feng Tian, Sicong Leng, Jiahao Nie 等CVPR 2025
- HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction DataQifan Yu, Juncheng Li, Longhui Wei, Liang Pang 等CVPR 2024 · 被引用 29 次
- See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMsZiyun Dai, Xiaoqiang Li, Shaohua Zhang, Yuanchen Wu 等ACM MM 2025
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
- Multi-Modal Hallucination Control by Visual Information GroundingAlessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary 等CVPR 2024
