Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Florian Boudin, Atsuhiro Takasu, Akiko Aizawa
Abstract
With the growing number of submitted scientific papers, there is an increasing demand for systems that can assist reviewers in evaluating research claims. Experimental results are a core component of scientific work, often presented in varying formats such as tables or charts. Understanding how robust current multimodal large language models (multimodal LLMs) are at verifying scientific claims across different evidence formats remains an important and underexplored challenge. In this paper, we design and conduct a series of experiments to assess the ability of multimodal LLMs to verify scientific claims using both tables and charts as evidence. To enable this evaluation, we adapt two existing datasets of scientific papers by incorporating annotations and structures necessary for a multimodal claim verification task. Using this adapted dataset, we evaluate 12 multimodal LLMs and find that current models perform better with table-based evidence while struggling with chart-based evidence. We further conduct human evaluations and observe that humans maintain strong performance across both formats, unlike the models. Our analysis also reveals that smaller multimodal LLMs (under 8B) show weak correlation in performance between table-based and chart-based tasks, indicating limited cross-modal generalization. These findings highlight a critical gap in current models' multimodal reasoning capabilities. We suggest that future multimodal LLMs should place greater emphasis on improving chart understanding to better support scientific claim verification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d5802ea-9237-430e-acd7-b881e8ade623Builds on9
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- SciVer: Evaluating Foundation Models for Multimodal Scientific Claim VerificationChengye Wang, Yifei Shen, Zexi Kuang, Arman Cohan et al.ACL 2025 · 8 citations
- SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific TablesXinyuan Lu, Liangming Pan, Qian Liu, Preslav Nakov et al.EMNLP 2023 · 7 citations
- Fact or Fiction: Verifying Scientific ClaimsDavid Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang et al.EMNLP 2020 · 6 citations
Related papers
- ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific ChartsRuiran Su, Jiasheng Si, Zhijiang Guo, Janet B. PierrehumbertEMNLP 2025
- mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language ModelAnwen Hu, Yaya Shi, Haiyang Xu, Jiabo Ye et al.ACM MM 2024 · 15 citations
- Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question AnsweringZixin Chen, Sicheng Song, KaShun Shum, Yanna Lin et al.EMNLP 2025 · 1 citation
- How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?Leo Yu-Ho Lo, Huamin QuIEEE VIS 2024 · 24 citations
- DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific ChartsYujing Lu, Ling Zhong, Jing Yang, Weiming Li et al.AAAI 2026
