Lune

ACL2026Top-tier venue

FactVerse: A Benchmark for Factual Consistency in Interleaved Image-Text Generation

Yubo Shan, Kun Zhang, Qiming Xu, Liping Cao, Yingying Cao, Jian Zhang, Yu Wang, Jingyuan Li, Yuanzhuo Wang

2026Year

Abstract

Interleaved multimodal understanding and generation-where models can interactively comprehend and produce images and text in arbitrary orders-has emerged as a key research direction in generative Multimodal Large Language Models(MLLMs). Such interleaved image-text content plays an increasingly important role in information dissemination. However, the compounded persuasive power of multimodal narratives also raises the risk of factual misinformation. Despite this, existing benchmarks lack effective mechanisms to evaluate factual consistency in interleaved image-text content. To bridge this gap, we introduce Fact-Verse, a benchmark dedicated to evaluating factual consistency in interleaved image-text generation. FactVerse comprises 3,000 humanverified instances across four categories and 50 domains, supporting both English and Chinese. We also establish a multi-dimensional evaluation framework designed to rigorously assess factual consistency. Experiments demonstrate that our framework achieves high alignment with human judgments, significantly outperforming existing evaluation methods. Furthermore, our analysis reveals systematic deficiencies in current models, offering critical insights for future design.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7e3f9e42-2155-424c-9be8-e21f5f193bc5

Builds on16

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines