Lune

NeurIPS2025顶会

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?

Hao Yin, Guangzong Si, Zilei Wang

2025年份
6被引次数
1顶会引用

摘要

Contrastive decoding strategies are widely used to reduce object hallucinations in multimodal large language models (MLLMs). These methods work by constructing contrastive samples to induce hallucinations and then suppressing them in the output distribution. However, this paper demonstrates that such approaches fail to effectively mitigate the hallucination problem. The performance improvements observed on POPE Benchmark are largely driven by two misleading factors: (1) crude, unidirectional adjustments to the model's output distribution and (2) the adaptive plausibility constraint, which reduces the sampling strategy to greedy search. To further illustrate these issues, we introduce a series of spurious improvement methods and evaluate their performance against contrastive decoding techniques. Experimental results reveal that the observed performance gains in contrastive decoding are entirely unrelated to its intended goal of mitigating hallucinations. Our findings challenge common assumptions about the effectiveness of contrastive decoding strategies and pave the way for developing genuinely effective solutions to hallucinations in MLLMs. The source code is available at https://github.com/ustc-hyin/cd_rethink * Corresponding Author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

methods fail to effectively address model hallucination. The observed performance gains on the POPE benchmark are primarily driven by two factors: Misleading Nature of Performance Improvement R 1 : A unidirectional adjustment of the output distribution, which simply biases the model towards producing more "Yes" outputs, leading to a balanced distribution on certain datasets.

R 2 : The adaptive constraints in these methods degrade the sampling decoding strategy into an approximation of greedy search, resulting in deceptively improved performance.

To expose the misleading nature of the improvement in the first scenario, we implemented two forced distribution adjustment algorithms in Section 5.1 to show that the apparent gains of contrastive decoding on the POPE Benchmark are not genuine. The methods are as follows: (1) Prompt-Based Adjustment, where we added a prompt to the instruction, such as "Whenever possible, please select Yes." to bias outputs toward "Yes"; and (2) Output Layer Modification, where we altered the output layer to favor "Yes" when the probabilities for "Yes" and "No" were similar. Although neither method mitigates hallucinations, both achieved performance gains comparable to those of contrastive decoding, confirming that these improvements do not represent a genuine solution to the problem.

To highlight the misleading nature of the performance improvement in the second scenario, we incorporated the adaptive plausibility constraint into the standard sampling strategy and compared its predictions with those from contrastive decoding in Section 5.2. The experimental results reveal that, despite having no theoretical connection to hallucination mitigation, the adaptive plausibility constraint accounts for nearly all the performance gains attributed to contrastive decoding. This finding underscores that the contrastive decoding methods, in essence, fail to mitigate hallucinations.

Overall, this paper makes the following three contributions: • We identified that the performance improvement of contrastive decoding methods stems from its unidirectional and blunt adjustment of the output distribution, which coincidentally balances the distribution on certain datasets.

• We discovered that another key factor driving the performance gains of contrastive decoding methods is their adaptive plausibility constraints, which streamline the sampling strategy into an approximation of greedy search.

• We developed a series of spurious improvement methods and evaluated their performance against contrastive decoding methods. Our findings convincingly show that contrastive decoding methods do not alleviate hallucinations in any meaningful way.

2 Related Work

Multimodal Large Language Models. The evolution of MLLMs [20,21] has progressed from BERTbased decoders [22,23] to advanced LLM architectures [24,25], enabling more effective multimodal relationship modeling [26,27]. Models such as BLIP-2 [28] and MiniGPT-4 [29] employ Q-Former mechanisms to enhance the alignment between visual and textual inputs, facilitating more precise cross-modal interactions. InstructBLIP [30] extends this framework by integrating task-specific instructions, improving the model's ability to interpret context-sensitive visual semantics. Meanwhile, LLaVA [31,32] and Qwen-VL [33] adopt simpler linear projection methods that streamline alignment, leading to superior performance in vision-language tasks. Despite these advancements, hallucination remains a persistent challenge that warrants further investigation.

Contrastive decoding [34,35,36] are widely recognized as effective in addressing object hallucination in generative mod

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖