Seeing Beyond Hallucinations: LLM-based Compositional Information Extraction for Multimodal Reasoning
Xinwei Li, Li Lin, Shuai Wang, Hanqian Wu
摘要
Advancements in Multimodal Large Language Models (MLLMs) have significantly improved information extraction and retrieval performance. Despite these achievements, MLLMs still suffer from the visual object hallucination problem, where models produce plausible, yet incorrect, or irrelevant content not present in the input data. This issue arises from an over-reliance on ''bag-of-objects'' representations and language priors, leading to inadequate extraction of visual objects, along with their attributes and relationships. Existing methods to mitigate these hallucinations are limited by the significant human labor required and the coarse-grained nature. To overcome these challenges, we introduce Multimodal Contrastive Decoding (MMCD), a novel decoding approach that integrates graph-structured reasoning paths with contrastive decoding. MMCD mitigates object hallucinations induced by language priors and enhances the ability of MLLMs to extract and understand compositional information, without additional training or the usage of external tools. This is achieved by masking key objects in images, constructing perturbed scene graphs of attributes and relationships, then contrasting these with the original image and scene graph. Extensive evaluation across three distinct multimodal compositional reasoning tasks: spatial relationship reasoning, alignment of synthetic image and caption, and fine-grained object attribute understanding, show that MMCD consistently surpasses existing decoding methods when applied to various MLLMs. Moreover, MMCD achieves state-of-the-art performance on multiple benchmarks, including the What's Up, SeeTrue and SugarCrepe datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question AnsweringZhiyang Li, Ao Ke, Yukun Cao, Xike XieACL 2026
- DiVE: Decoupling Intra-layer Visual Evidence for Mitigating Hallucinations in Large Vision-Language ModelsXinwei Li, Li Lin, Hui Jiao, Li Yao 等ACL 2026
相关 Paper
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive DecodingSicong Leng, Hang Zhang, Guanzheng Chen, Xin Li 等CVPR 2024
- Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language ModelsWei Chen, Xin Yan, Bin Wen, Fan Yang 等NeurIPS 2025 · 被引用 4 次
- Multi-Frequency Contrastive Decoding: Alleviating Hallucinations for Large Vision-Language ModelsBingqian Liu, Fu Zhang, Guoqing Chen, Jingwei ChengEMNLP 2025
- CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal ModelsJunho Kim, Hyunjun Kim, Yeonju Kim, Yong Man RoNeurIPS 2024 · 被引用 55 次
- VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language ModelsGuoqing Chen, Fu Zhang, Bingqian Liu, Chenglong Lu 等AAAI 2026
