RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
Xuming He, Zhiyuan You, Junchao Gong, Couhua Liu, Xiaoyu Yue, Peiqin Zhuang, Wenlong Zhang, Lei Bai
Abstract
Quality analysis of weather forecasts is an essential topic in meteorology. Although traditional score-based evaluation metrics can quantify certain forecast errors, they are still far from meteorological experts in terms of descriptive capability, interpretability, and understanding of dynamic evolution. With the rapid development of Multi-modal Large Language Models (MLLMs), these models become potential tools to overcome the above challenges. In this work, we introduce an MLLM-based weather forecast analysis method, RadarQA, integrating key physical attributes with detailed assessment reports. We introduce a novel and comprehensive task paradigm for multi-modal quality analysis, encompassing both single frame and sequence, under both rating and assessment scenarios. To support training and benchmarking, we design a hybrid annotation pipeline that combines human expert labeling with automated heuristics. With such an annotation method, we construct RQA-70K, a large-scale dataset with varying difficulty levels for radar forecast quality evaluation. We further design a multi-stage training strategy that iteratively improves model performance at each stage. Extensive experiments show that RadarQA outperforms existing general MLLMs across all evaluation settings, highlighting its potential for advancing quality analysis in weather prediction. The code and dataset are publicly available at https://github.com/hexmSeeU/RadarQA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d20684fb-3ba8-4666-8291-e95277a95ce2Cited by top-tier papers2
- Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and GenerationZhiwang Zhou, Yuandong Pu, Xuming He, Yidi Liu et al.ICLR 2026
- SynWeather: Weather Observation Data Synthesis Across Multiple Regions and Variables via a General Diffusion TransformerKaiyi Xu, Junchao Gong, Zhiwang Zhou, Zhangrui Li et al.AAAI 2026
Builds on22
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
Related papers
- Revisiting MLLM Based Image Quality Assessment: Errors and RemedyZhenchen Tang, Songlin Yang, Bo Peng, Zichuan Wang et al.AAAI 2026 · 2 citations
- WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report GenerationZinan Zheng, Yang Liu, Nuo Chen, Juepeng Zheng et al.ICML 2026
- ClimaQA: An Automated Evaluation Framework for Climate Question Answering ModelsVeeramakali Vignesh Manivannan, Yasaman Jafari, Srikar Eranky, Spencer Ho et al.ICLR 2025
- RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge InjectionWenjun Hou, Yi Cheng, Kaishuai Xu, Heng Li et al.ACL 2025
- MMClima: A Framework for Multimodal Climate Science Data and EvaluationMuhammad Umer Sheikh, Hassan Abid, Khawar shehzad, Ufaq Khan et al.ICML 2026 · 2 citations
