Do Large Language Models Reason About Uncertainty Like Humans? A Benchmark on Hurricane Forecast Visualization Comprehension
Le Liu, Yuhao Wang, Bohan Shen, Wei Zeng, Shizhou Zhang, Di Xu, Peng Wang
摘要
Uncertainty visualizations, such as hurricane cones and ensemble tracks, are essential for risk communication but are often misinterpreted, leading to harmful decisions. As AI assistants like large language models (LLMs) increasingly support understanding of graphics and decision-making, they offer a promising pathway to enhance the interpretation of complex visualizations and a new opportunity to examine and improve the interpretation of uncertainty. We introduce UnReason, the first benchmark that systematically compares how humans and LLMs reason about hurricane forecast uncertainty visualizations. UnReason spans two escalating phases, seven representative visualization formats, six real hurricane cases, and three agent types (humans, LLMs with context, and LLMs without context), including 880 visualizations and 117,600 structured question–answer pairs under matched evaluation conditions. Phase 1 evaluates reasoning across implicit and explicit uncertainty encodings; Phase 2 examines reasoning under single- versus multi-dimensional uncertainty representations. We thoroughly assess damage estimation, reasoning strategies, and comprehension patterns, revealing that LLMs have a stronger semantic and conceptual understanding of uncertainty, and are less misled by visual variability, but still replicate key human biases during decision-making. Our findings offer insights into aligning LLM behavior with human cognition in uncertainty-rich visual reasoning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- An Empirical Evaluation of the GPT-4 Multimodal Language Model on Visualization Literacy TasksAlexander Bendeck, John T. StaskoIEEE VIS 2024 · 被引用 40 次
- WorldGPT: Empowering LLM as Multimodal World ModelZhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li 等ACM MM 2024 · 被引用 35 次
- Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction TuningXingchen Zeng, Haichuan Lin, Yilin Ye, Wei ZengIEEE VIS 2024 · 被引用 23 次
- Tree-of-Reasoning Question Decomposition for Complex Question Answering with Large Language ModelsKun Zhang, Jiali Zeng, Fandong Meng, Yuanzhuo Wang 等AAAI 2024 · 被引用 14 次
相关 Paper
- ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question AnsweringXiaojun Chen, Sixiao Luo, Ziqi Liu, Min Yang 等CVPR 2026
- InsLogicBench: An Argumentation Logic Grounded Benchmark for Complex Insurance Claims AdjudicationJin Liu, Yunpeng Liu, Keyi Wang, Jie Shi 等ACL 2026
- Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question AnsweringZixin Chen, Sicheng Song, KaShun Shum, Yanna Lin 等EMNLP 2025 · 被引用 1 次
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen 等ICLR 2026 · 被引用 103 次
- DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense DocumentsZhuoran Yu, Le T Nguyen, Jaden Park, Xinyi Gu 等ICML 2026
