Seeing Beyond Hallucinations: LLM-based Compositional Information Extraction for Multimodal Reasoning
Xinwei Li, Li Lin, Shuai Wang, Hanqian Wu
Abstract
Advancements in Multimodal Large Language Models (MLLMs) have significantly improved information extraction and retrieval performance. Despite these achievements, MLLMs still suffer from the visual object hallucination problem, where models produce plausible, yet incorrect, or irrelevant content not present in the input data. This issue arises from an over-reliance on ''bag-of-objects'' representations and language priors, leading to inadequate extraction of visual objects, along with their attributes and relationships. Existing methods to mitigate these hallucinations are limited by the significant human labor required and the coarse-grained nature. To overcome these challenges, we introduce Multimodal Contrastive Decoding (MMCD), a novel decoding approach that integrates graph-structured reasoning paths with contrastive decoding. MMCD mitigates object hallucinations induced by language priors and enhances the ability of MLLMs to extract and understand compositional information, without additional training or the usage of external tools. This is achieved by masking key objects in images, constructing perturbed scene graphs of attributes and relationships, then contrasting these with the original image and scene graph. Extensive evaluation across three distinct multimodal compositional reasoning tasks: spatial relationship reasoning, alignment of synthetic image and caption, and fine-grained object attribute understanding, show that MMCD consistently surpasses existing decoding methods when applied to various MLLMs. Moreover, MMCD achieves state-of-the-art performance on multiple benchmarks, including the What's Up, SeeTrue and SugarCrepe datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a3778f81-38a5-4491-82ee-e7228c24434cCited by top-tier papers2
- KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question AnsweringZhiyang Li, Ao Ke, Yukun Cao, Xike XieACL 2026
- DiVE: Decoupling Intra-layer Visual Evidence for Mitigating Hallucinations in Large Vision-Language ModelsXinwei Li, Li Lin, Hui Jiao, Li Yao et al.ACL 2026
Related papers
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive DecodingSicong Leng, Hang Zhang, Guanzheng Chen, Xin Li et al.CVPR 2024
- Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language ModelsWei Chen, Xin Yan, Bin Wen, Fan Yang et al.NeurIPS 2025 · 4 citations
- Multi-Frequency Contrastive Decoding: Alleviating Hallucinations for Large Vision-Language ModelsBingqian Liu, Fu Zhang, Guoqing Chen, Jingwei ChengEMNLP 2025
- CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal ModelsJunho Kim, Hyunjun Kim, Yeonju Kim, Yong Man RoNeurIPS 2024 · 55 citations
- VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language ModelsGuoqing Chen, Fu Zhang, Bingqian Liu, Chenglong Lu et al.AAAI 2026
