OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, Nenghai Yu
摘要
Hallucination, posed as a pervasive challenge of multimodal large language models (MLLMs), has significantly impeded their real-world usage that demands precise judgment. Existing methods mitigate this issue with either training with specific designed data or inferencing with external knowledge from other sources, incurring inevitable additional costs. In this paper, we present OPERA, a novel MLLM decoding method grounded in an Over-trust Penalty and a Retrospection-Allocation strategy, serving as a nearly free lunch to alleviate the hallucination issue without additional data, knowledge, or training. Our approach begins with an interesting observation that, most hallucinations are closely tied to the knowledge aggregation patterns manifested in the self-attention matrix, i.e., MLLMs tend to generate new tokens by focusing on a few summary tokens, but not all the previous tokens. Such partial overtrust inclination results in the neglecting of image tokens and describes the image content with hallucination. Based on the observation, OPERA introduces a penalty term on the model logits during the beam-search decoding to mitigate the over-trust issue, along with a rollback strategy that retrospects the presence of summary tokens in the previously generated tokens, and re-allocate the token selection if necessary. With extensive experiments, OPERA shows significant hallucination-mitigating performance on different MLLMs and metrics, proving its effectiveness and generality. Our code is available at: This link.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper169
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning ModelsZhongxing Xu, Chengzhi Liu, Qingyue Wei, Juncheng Wu 等NeurIPS 2025 · 被引用 103 次
- VIGC: Visual Instruction Generation and CorrectionBin Wang, Fan Wu, Xiao Han, Jiahui Peng 等AAAI 2024 · 被引用 95 次
- Latent Visual ReasoningBangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang 等ICLR 2026 · 被引用 80 次
- Mitigating Object Hallucination via Concentric Causal AttentionYun Xing, Yiheng Li, Ivan Laptev, Shijian LuNeurIPS 2024 · 被引用 78 次
- Multi-Object Hallucination in Vision Language ModelsXuweiyi Chen, Ziqiao Ma, Xuejun Zhang, Sihan Xu 等NeurIPS 2024 · 被引用 77 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- DOPRA: Decoding Over-accumulation Penalization and Re-allocation in Specific Weighting LayerJinfeng Wei, Xiaofeng ZhangACM MM 2024 · 被引用 25 次
- MLLM can see? Dynamic Correction Decoding for Hallucination MitigationChenxi Wang, Xiang Chen, Ningyu Zhang, Bozhong Tian 等ICLR 2025
- Intervening Anchor Token: Decoding Strategy in Alleviating Hallucinations for MLLMsFeilong Tang, Zile Huang, Chengzhi Liu, Qiang Sun 等ICLR 2025
- Attributive Reasoning for Hallucination Diagnosis of Large Language ModelsYuyan Chen, Zehao Li, Shuangjie You, Zhengyu Chen 等AAAI 2025 · 被引用 33 次
- Controlling Multimodal Llms Via Reward-Guided DecodingOscar Mañas, Pierluca D'Oro, Koustuv Sinha, Adriana Romero-Soriano 等ICCV 2025
