Lune

EMNLP2025顶会

ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts

Ruiran Su, Jiasheng Si, Zhijiang Guo, Janet B. Pierrehumbert

2025年份
1顶会引用

摘要

Scientific fact-checking has largely focused on textual and tabular sources, neglecting scientific charts-a primary medium for conveying quantitative evidence and supporting statistical reasoning in research communication. We introduce CLIMATEVIZ, the first largescale benchmark for scientific fact-checking grounded in real-world, expert-curated scientific charts. CLIMATEVIZ comprises 49,862 claims paired with 2,896 visualizations, each labeled as support, refute, or not enough information. To enable interpretable verification, each instance includes structured knowledge graph explanations that capture statistical patterns, temporal trends, spatial comparisons, and causal relations. We conduct a comprehensive evaluation of state-of-the-art multimodal large language models, including proprietary and open-source systems, under zero-shot and few-shot settings. Our results show that current models struggle to perform fact-checking when statistical reasoning over charts is required: even the best-performing systems, such as Gemini 2.5 and InternVL 2.5, achieve only 76.2-77.8% accuracy in label-only output settings, which is far below human performance (89.3% and 92.7%). While few-shot prompting yields limited improvements, explanationaugmented outputs significantly enhance performance in some closed-source models, notably o3 and Gemini 2.5. We released our dataset and code alongside the paper. 1 (c) Subgraph of Relevant Facts Caption: Cumulative mass loss of the Greenland Ice Sheet from 1972 to 2022, showing accelerating ice loss and corresponding sea level rise.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖