Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
Angela van Sprang, Laurens Samson, Ana Lucic, Erman Acar, Sennay Ghebreab, Yuki M Asano
Abstract
Multimodal large language models (MLLMs) are trained to represent vision and language in a shared space. But does this joint representation enable consistent reasoning across modalities? We introduce REST and REST+ (Render-Equivalence Stress Tests), two benchmarks for systematically evaluating cross-modal consistency. Each sample presents semantically identical information in three forms (image, text, and mixed), allowing us to measure whether models produce consistent outputs regardless of modality. Evaluating 15 state-of-the-art MLLMs, we find that none reason consistently across modalities, with substantial variation in the degree of inconsistency. Neither rendering text as images nor images as text resolves this problem, even when controlling for OCR errors. We further show that visual characteristics (color, resolution, but not font) and the number of vision tokens affect performance even when text is correctly recognized. Finally, our consistency score correlates with the cross-modal cosine similarity in embedding space, suggesting a mechanistic explanation: inconsistent reasoning arises when text and image representations occupy distinct regions of the joint space. Data and code are available on https://github.com/ angelavansprang/Same-Content-Different- Answers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e27aa07e-0afb-4d02-9e7b-2e53e4c92d35Cited by top-tier papers2
- Mechanisms of Prompt-Induced Hallucination in Vision-Language ModelsWilliam Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov et al.ACL 2026 · 4 citations
- Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-ExpertsHaolei Xu, Haiwen Hong, Hongxing Li, Rui Zhou et al.ACL 2026 · 3 citations
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXingrui Wang, Jiang Liu, Chao Huang, Xiaodong Yu et al.ICLR 2026 · 4 citations
- Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response TheoryShunki Uebayashi, Kento Masui, Kyohei Atarashi, Han Bao et al.ICLR 2026 · 1 citation
- Unveiling the Tapestry of Consistency in Large Vision-Language ModelsYuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan et al.NeurIPS 2024 · 27 citations
- PRISM: A Benchmark for Unveiling Cross-modal Knowledge Inconsistency in Large Vision-Language ModelsMingjie Wei, Wei-Nan Zhang, Chen Zhang, Yifeng Ding et al.ACM MM 2025 · 1 citation
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image ReasoningMingxin Huang, Yongxin Shi, Dezhi Peng, Songxuan Lai et al.ICLR 2026 · 28 citations
