ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts
Ruiran Su, Jiasheng Si, Zhijiang Guo, Janet B. Pierrehumbert
摘要
Scientific fact-checking has largely focused on textual and tabular sources, neglecting scientific charts-a primary medium for conveying quantitative evidence and supporting statistical reasoning in research communication. We introduce CLIMATEVIZ, the first largescale benchmark for scientific fact-checking grounded in real-world, expert-curated scientific charts. CLIMATEVIZ comprises 49,862 claims paired with 2,896 visualizations, each labeled as support, refute, or not enough information. To enable interpretable verification, each instance includes structured knowledge graph explanations that capture statistical patterns, temporal trends, spatial comparisons, and causal relations. We conduct a comprehensive evaluation of state-of-the-art multimodal large language models, including proprietary and open-source systems, under zero-shot and few-shot settings. Our results show that current models struggle to perform fact-checking when statistical reasoning over charts is required: even the best-performing systems, such as Gemini 2.5 and InternVL 2.5, achieve only 76.2-77.8% accuracy in label-only output settings, which is far below human performance (89.3% and 92.7%). While few-shot prompting yields limited improvements, explanationaugmented outputs significantly enhance performance in some closed-source models, notably o3 and Gemini 2.5. We released our dataset and code alongside the paper. 1 (c) Subgraph of Relevant Facts Caption: Cumulative mass loss of the Greenland Ice Sheet from 1972 to 2022, showing accelerating ice loss and corresponding sea level rise.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
相关 Paper
- Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question AnsweringZixin Chen, Sicheng Song, KaShun Shum, Yanna Lin 等EMNLP 2025 · 被引用 1 次
- DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific ChartsYujing Lu, Ling Zhong, Jing Yang, Weiming Li 等AAAI 2026
- RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo 等ACL 2026
- Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data ExtractionAmit Kumar Das, Mohammad Tarun, Klaus MuellerIEEE VIS 2025 · 被引用 6 次
- MMClima: A Framework for Multimodal Climate Science Data and EvaluationMuhammad Umer Sheikh, Hassan Abid, Khawar shehzad, Ufaq Khan 等ICML 2026 · 被引用 2 次
