CLLMate: A Multimodal Benchmark for Weather and Climate Events Forecasting
Haobo Li, Zhaowei Wang, Jiachen Wang, Yueya Wang, Alexis Kai-Hon Lau, Huamin Qu
Abstract
Forecasting weather and climate events is crucial for making appropriate measures to mitigate environmental hazards and minimize losses. However, existing environmental forecasting research focuses narrowly on predicting numerical meteorological variables (e.g., temperature), neglecting the translation of these variables into actionable textual narratives of events and their consequences. To bridge this gap, we proposed Weather and Climate Event Forecasting (WCEF), a new task that leverages numerical meteorological raster data and textual event data to predict weather and climate events. This task is challenging to accomplish due to difficulties in aligning multimodal data and the lack of supervised datasets. To address these challenges, we present CLLMate, the first multimodal dataset for WCEF, using 26,156 environmental news articles aligned with ERA5 reanalysis data. We systematically benchmark 32 existing models on CLLMate, including closed-source, open-source, and our fine-tuned models. Our experiments reveal the advantages and limitations of existing MLLMs and the value of CLLMate for the training and benchmarking of the WCEF task. The dataset is available at https://github.com/hobolee/ CLLMate .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- MM-Forecast: A Multimodal Approach to Temporal Event Forecasting with Large Language ModelsHaoxuan Li, Zhengmao Yang, Yunshan Ma, Yi Bin et al.ACM MM 2024 · 4 citations
- ForecastQA: A Question Answering Challenge for Event Forecasting with Temporal Text DataWoojeong Jin, Rahul Khanna, Suji Kim, Dong-Ho Lee et al.ACL 2021
- CrisisTS: Coupling Social Media Textual Data and Meteorological Time Series for Urgency ClassificationRomain Meunier, Farah Benamara, Véronique Moriceau, Zhongzheng Qiao et al.ACL 2025 · 1 citation
- WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report GenerationZinan Zheng, Yang Liu, Nuo Chen, Juepeng Zheng et al.ICML 2026
- Image Enhanced Event Detection in News ArticlesMeihan Tong, Shuai Wang, Yixin Cao, Bin Xu et al.AAAI 2020 · 43 citations
