PRISM: A Benchmark for Unveiling Cross-modal Knowledge Inconsistency in Large Vision-Language Models
Mingjie Wei, Wei-Nan Zhang, Chen Zhang, Yifeng Ding, Donglin Di, Lei Ren, Wei Chen, Ting Liu
Abstract
Recent advances in Large Vision-Language Models (LVLMs) have unearthed boosted performance of multi-modal understanding. In this paper, however, we for the first time uncover a critically under-explored challenge persisting in this trend, that LVLMs unfortunately exhibit cross-modal knowledge inconsistencies. Cross-modal knowledge inconsistency refers to the tendency of providing semantically inconsistent responses to contexts that are semantically equivalent but expressed in different modalities. In real-world applications, users can rely on either text or image to express their ideas. Inconsistent responses across modalities can confuse users, challenging the reliabilities of LVLMs in practice. Therefore, we argue that evaluating performance on either multi-modal or text-only task is insufficient; and waiving the mentioned cross-modal knowledge inconsistency is crucial. The paper proposes PRISM, the first-ever benchmark for measuring the inconsistency, and the corresponding evaluation metric Know-Inc. PRISM covers commonsense, encyclopedia, and mathematics knowledge, with manually-screened samples of semantic alignment. From the evaluation results of up to 27 LVLMs with diverse structures, we conclude that: 1) LVLMs show a preference for textual input, 2) there is a correlation between inconsistency and accuracy, and 3) the inconsistency is more prominent in encyclopedia knowledge. These findings can shed light on further optimization and development of LVLMs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fefa3f7a-8016-4eae-901f-ff07aa4ddd37Cited by top-tier papers1
Ask how each one uses itRelated papers
- Unveiling the Tapestry of Consistency in Large Vision-Language ModelsYuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan et al.NeurIPS 2024 · 27 citations
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal InconsistenciesLukas Selch, Yufang Hou, Muhammad Jehanzeb Mirza, Sivan Doveh et al.ICLR 2026 · 2 citations
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXingrui Wang, Jiang Liu, Chao Huang, Xiaodong Yu et al.ICLR 2026 · 4 citations
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict ResolutionBaoliang Tian, Yuxuan Si, Jilong Wang, Lingyao Li et al.AAAI 2026 · 2 citations
- Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMsXiaoyuan Liu, Wenxuan Wang, Youliang Yuan, Jen-tse Huang et al.ACL 2025 · 20 citations
