Lune

CVPR2026Top-tier venue

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs

Angela van Sprang, Laurens Samson, Ana Lucic, Erman Acar, Sennay Ghebreab, Yuki M Asano

2026Year
8Citations
2Top-tier citations

Abstract

Multimodal large language models (MLLMs) are trained to represent vision and language in a shared space. But does this joint representation enable consistent reasoning across modalities? We introduce REST and REST+ (Render-Equivalence Stress Tests), two benchmarks for systematically evaluating cross-modal consistency. Each sample presents semantically identical information in three forms (image, text, and mixed), allowing us to measure whether models produce consistent outputs regardless of modality. Evaluating 15 state-of-the-art MLLMs, we find that none reason consistently across modalities, with substantial variation in the degree of inconsistency. Neither rendering text as images nor images as text resolves this problem, even when controlling for OCR errors. We further show that visual characteristics (color, resolution, but not font) and the number of vision tokens affect performance even when text is correctly recognized. Finally, our consistency score correlates with the cross-modal cosine similarity in embedding space, suggesting a mechanistic explanation: inconsistent reasoning arises when text and image representations occupy distinct regions of the joint space. Data and code are available on https://github.com/ angelavansprang/Same-Content-Different- Answers.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e27aa07e-0afb-4d02-9e7b-2e53e4c92d35

Cited by top-tier papers2

Ask how each one uses it

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines