Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
Xiaoyuan Liu, Wenxuan Wang, Youliang Yuan, Jen-tse Huang, Qiuzhi Liu, Pinjia He, Zhaopeng Tu
摘要
This paper explores the problem of commonsense level vision-knowledge conflict in Multimodal Large Language Models (MLLMs), where visual information contradicts model's internal commonsense knowledge. To study this issue, we introduce an automated framework, augmented with human-in-the-loop quality control, to generate inputs designed to simulate and evaluate these conflicts in MLLMs. Using this framework, we have crafted a diagnostic benchmark consisting of 374 original images and 1,122 high-quality question-answer (QA) pairs. The benchmark covers two aspects of conflict and three question types, providing a thorough assessment tool. We apply this benchmark to assess the conflict-resolution capabilities of nine representative MLLMs from various model families. Our results indicate an evident over-reliance on parametric knowledge for approximately 20% of all queries, especially among Yes-No and action-related problems. Based on these findings, we evaluate the effectiveness of existing approaches to mitigating the conflicts and compare them to our "Focuson-Vision" prompting strategy. Despite some improvement, the vision-knowledge conflict remains unresolved and can be further scaled through our data construction framework. Our proposed framework, benchmark, and analysis contribute to the understanding and mitigation of vision-knowledge conflicts in MLLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language ModelsHengzhuang Li, Xinsong Zhang, QIMING PENG, Bin Luo 等CVPR 2026 · 被引用 2 次
- REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot AlignmentKai Ye, Xianwei Mao, Sheng Zhou, Zirui Shao 等ICML 2026 · 被引用 1 次
- Fix Before Search: Benchmarking Agentic Visual Query Pre-processing in Multimodal Retrieval-augmented GenerationShenglai Zeng, Jiankun Zhang, Kai Guo, Xinnan Dai 等ICML 2026 · 被引用 1 次
- When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMsBeidi Zhao, wenlong deng, Xinting Liao, Yushu Li 等ICML 2026 · 被引用 1 次
- SHARP: Steering Hallucination in LVLMs via Representation EngineeringJunfei Wu, Yue Ding, Guofan Liu, Tianze Xia 等EMNLP 2025
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- Unveiling the Tapestry of Consistency in Large Vision-Language ModelsYuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan 等NeurIPS 2024 · 被引用 27 次
- VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsMingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li 等AAAI 2026
- PRISM: A Benchmark for Unveiling Cross-modal Knowledge Inconsistency in Large Vision-Language ModelsMingjie Wei, Wei-Nan Zhang, Chen Zhang, Yifeng Ding 等ACM MM 2025 · 被引用 1 次
- Benchmarking Deflection and Hallucination in Large Vision-Language ModelsNicholas Moratelli, Christopher Davis, Leonardo F. R. Ribeiro, Bill Byrne 等ACL 2026 · 被引用 1 次
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict ResolutionBaoliang Tian, Yuxuan Si, Jilong Wang, Lingyao Li 等AAAI 2026 · 被引用 2 次
