Lune

NeurIPS2025Top-tier venue

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?

Hao Yin, Guangzong Si, Zilei Wang

2025Year
6Citations
1Top-tier citations

Abstract

Contrastive decoding strategies are widely used to reduce object hallucinations in multimodal large language models (MLLMs). These methods work by constructing contrastive samples to induce hallucinations and then suppressing them in the output distribution. However, this paper demonstrates that such approaches fail to effectively mitigate the hallucination problem. The performance improvements observed on POPE Benchmark are largely driven by two misleading factors: (1) crude, unidirectional adjustments to the model's output distribution and (2) the adaptive plausibility constraint, which reduces the sampling strategy to greedy search. To further illustrate these issues, we introduce a series of spurious improvement methods and evaluate their performance against contrastive decoding techniques. Experimental results reveal that the observed performance gains in contrastive decoding are entirely unrelated to its intended goal of mitigating hallucinations. Our findings challenge common assumptions about the effectiveness of contrastive decoding strategies and pave the way for developing genuinely effective solutions to hallucinations in MLLMs. The source code is available at https://github.com/ustc-hyin/cd_rethink * Corresponding Author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

methods fail to effectively address model hallucination. The observed performance gains on the POPE benchmark are primarily driven by two factors: Misleading Nature of Performance Improvement R 1 : A unidirectional adjustment of the output distribution, which simply biases the model towards producing more "Yes" outputs, leading to a balanced distribution on certain datasets.

R 2 : The adaptive constraints in these methods degrade the sampling decoding strategy into an approximation of greedy search, resulting in deceptively improved performance.

To expose the misleading nature of the improvement in the first scenario, we implemented two forced distribution adjustment algorithms in Section 5.1 to show that the apparent gains of contrastive decoding on the POPE Benchmark are not genuine. The methods are as follows: (1) Prompt-Based Adjustment, where we added a prompt to the instruction, such as "Whenever possible, please select Yes." to bias outputs toward "Yes"; and (2) Output Layer Modification, where we altered the output layer to favor "Yes" when the probabilities for "Yes" and "No" were similar. Although neither method mitigates hallucinations, both achieved performance gains comparable to those of contrastive decoding, confirming that these improvements do not represent a genuine solution to the problem.

To highlight the misleading nature of the performance improvement in the second scenario, we incorporated the adaptive plausibility constraint into the standard sampling strategy and compared its predictions with those from contrastive decoding in Section 5.2. The experimental results reveal that, despite having no theoretical connection to hallucination mitigation, the adaptive plausibility constraint accounts for nearly all the performance gains attributed to contrastive decoding. This finding underscores that the contrastive decoding methods, in essence, fail to mitigate hallucinations.

Overall, this paper makes the following three contributions: • We identified that the performance improvement of contrastive decoding methods stems from its unidirectional and blunt adjustment of the output distribution, which coincidentally balances the distribution on certain datasets.

• We discovered that another key factor driving the performance gains of contrastive decoding methods is their adaptive plausibility constraints, which streamline the sampling strategy into an approximation of greedy search.

• We developed a series of spurious improvement methods and evaluated their performance against contrastive decoding methods. Our findings convincingly show that contrastive decoding methods do not alleviate hallucinations in any meaningful way.

2 Related Work

Multimodal Large Language Models. The evolution of MLLMs [20,21] has progressed from BERTbased decoders [22,23] to advanced LLM architectures [24,25], enabling more effective multimodal relationship modeling [26,27]. Models such as BLIP-2 [28] and MiniGPT-4 [29] employ Q-Former mechanisms to enhance the alignment between visual and textual inputs, facilitating more precise cross-modal interactions. InstructBLIP [30] extends this framework by integrating task-specific instructions, improving the model's ability to interpret context-sensitive visual semantics. Meanwhile, LLaVA [31,32] and Qwen-VL [33] adopt simpler linear projection methods that streamline alignment, leading to superior performance in vision-language tasks. Despite these advancements, hallucination remains a persistent challenge that warrants further investigation.

Contrastive decoding [34,35,36] are widely recognized as effective in addressing object hallucination in generative mod

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f0e2a4a8-fb99-4913-ac89-305586eac08d

Cited by top-tier papers1

Ask how each one uses it

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines