ClimateSOM: A Visual Analysis Workflow for Climate Ensemble Datasets
Yuya Kawakami, Daniel Cayan, Dongyu Liu, Kwan-Liu Ma
Abstract
Ensemble datasets are ever more prevalent in various scientific domains. In climate science, ensemble datasets are used to capture variability in projections under plausible future conditions including greenhouse and aerosol emissions. Each ensemble model run produces projections that are fundamentally similar yet meaningfully distinct. Understanding this variability among ensemble model runs and analyzing its magnitude and patterns is a vital task for climate scientists. In this paper, we present ClimateSOM, a visual analysis workflow that leverages a self-organizing map (SOM) and Large Language Models (LLMs) to support interactive exploration and interpretation of climate ensemble datasets. The workflow abstracts climate ensemble model runs-spatiotemporal time series-into a distribution over a 2D space that captures the variability among the ensemble model runs using a SOM. LLMs are integrated to assist in sensemaking of this SOM-defined 2D space, the basis for the visual analysis tasks. In all, ClimateSOM enables users to explore the variability among ensemble model runs, identify patterns, compare and cluster the ensemble model runs. To demonstrate the utility of ClimateSOM, we apply the workflow to an ensemble dataset of precipitation projections over California and the Northwestern United States. Furthermore, we conduct a short evaluation of our LLM integration, and conduct an expert review of the visual workflow and the insights from the case studies with six domain experts to evaluate our approach and its utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- The Dawn of Natural Language to SQL: Are We Fully Ready? [Experiment, Analysis & Benchmark ]Boyan Li, Yuyu Luo, Chengliang Chai, Guoliang Li et al.VLDB 2024 · 137 citations
- Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language ModelsMichael Xieyang Liu, Tongshuang Wu, Tianying Chen, Franklin Mingzhe Li et al.CHI 2024 · 44 citations
- IRVINE: A Design Study on Analyzing Correlation Patterns of Electrical EnginesJoscha Eirich, Jakob Bonart, Dominik Jäckle, Michael Sedlmair et al.IEEE VIS 2021 · 37 citations
- A Qualitative Analysis of Common Practices in Annotations: A Taxonomy and Design SpaceMd Dilshadur Rahman, Ghulam Jilani Quadri, Bhavana Doppalapudi, Danielle Albers Szafir et al.IEEE VIS 2024 · 21 citations
Related papers
- Can LLMs Bridge Domain and Visualization? A Case Study on High-Dimension Data Visualization in Single-Cell TranscriptomicsQianwen Wang, Xinyi Liu, Nils GehlenborgIEEE VIS 2025 · 1 citation
- Thematic-LM: A LLM-based Multi-agent System for Large-scale Thematic AnalysisTingrui Qiao, Caroline Walker, Chris Cunningham, Yun Sing KohWWW 2025 · 32 citations
- GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision SupportMuhammad Umer Sheikh, Khawar Shehzad, Salman Khan, Fahad Shahbaz Khan et al.ACL 2026 · 1 citation
- How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying LayoutsHuichen Will Wang, Jane Hoffswell, Sao Myat Thazin Thane, Victor S. Bursztyn et al.IEEE VIS 2024 · 12 citations
- ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language ModelsHaoxuan Li, Zhen Wen, Qiqi Jiang, Chenxiao Li et al.IEEE VIS 2025 · 3 citations
