Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
Angela van Sprang, Laurens Samson, Ana Lucic, Erman Acar, Sennay Ghebreab, Yuki M Asano
摘要
Multimodal large language models (MLLMs) are trained to represent vision and language in a shared space. But does this joint representation enable consistent reasoning across modalities? We introduce REST and REST+ (Render-Equivalence Stress Tests), two benchmarks for systematically evaluating cross-modal consistency. Each sample presents semantically identical information in three forms (image, text, and mixed), allowing us to measure whether models produce consistent outputs regardless of modality. Evaluating 15 state-of-the-art MLLMs, we find that none reason consistently across modalities, with substantial variation in the degree of inconsistency. Neither rendering text as images nor images as text resolves this problem, even when controlling for OCR errors. We further show that visual characteristics (color, resolution, but not font) and the number of vision tokens affect performance even when text is correctly recognized. Finally, our consistency score correlates with the cross-modal cosine similarity in embedding space, suggesting a mechanistic explanation: inconsistent reasoning arises when text and image representations occupy distinct regions of the joint space. Data and code are available on https://github.com/ angelavansprang/Same-Content-Different- Answers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Mechanisms of Prompt-Induced Hallucination in Vision-Language ModelsWilliam Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov 等ACL 2026 · 被引用 4 次
- Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-ExpertsHaolei Xu, Haiwen Hong, Hongxing Li, Rui Zhou 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXingrui Wang, Jiang Liu, Chao Huang, Xiaodong Yu 等ICLR 2026 · 被引用 4 次
- Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response TheoryShunki Uebayashi, Kento Masui, Kyohei Atarashi, Han Bao 等ICLR 2026 · 被引用 1 次
- Unveiling the Tapestry of Consistency in Large Vision-Language ModelsYuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan 等NeurIPS 2024 · 被引用 27 次
- PRISM: A Benchmark for Unveiling Cross-modal Knowledge Inconsistency in Large Vision-Language ModelsMingjie Wei, Wei-Nan Zhang, Chen Zhang, Yifeng Ding 等ACM MM 2025 · 被引用 1 次
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image ReasoningMingxin Huang, Yongxin Shi, Dezhi Peng, Songxuan Lai 等ICLR 2026 · 被引用 28 次
