Scaling Laws of Global Weather Models
Yuejiang Yu, Langwen Huang, Alexandru Calotoiu, Torsten Hoefler
Abstract
Data-driven models are revolutionizing weather forecasting. To optimize training efficiency and model performance, this paper analyzes empirical scaling laws within this domain. We investigate the relationship between model performance (validation loss) and three key factors: model size (), dataset size (), and compute budget (). Across a range of models, we find that Aurora exhibits the strongest data-scaling behavior: increasing the training dataset by 10× reduces validation loss by up to 3.2×. GraphCast demonstrates the highest parameter efficiency, yet suffers from limited hardware utilization. Our compute-optimal analysis indicates that, under fixed compute budgets, allocating resources to more total training data yields greater performance gains than increasing model size. Furthermore, we analyze model shape and uncover scaling behaviors that differ fundamentally from those observed in language models: weather forecasting models consistently favor increased width over depth. These findings suggest that future weather models should prioritize wider architectures and larger effective training datasets to maximize predictive performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dece11f1-2a96-4d66-8ca9-be8303a2d8f9Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Spherical Fourier Neural Operators: Learning Stable Dynamics on the SphereBoris Bonev, Thorsten Kurth, Christian Hundt, Jaideep Pathak et al.ICML 2023 · 280 citations
- Scaling transformer neural networks for skillful and reliable medium-range weather forecastingTung Nguyen, Rohan Shah, Hritik Bansal, Troy Arcomano et al.NeurIPS 2024 · 165 citations
- Scaling Laws for a Multi-Agent Reinforcement Learning ModelOren Neumann, Claudius GrosICLR 2023 · 3 citations
Related papers
- EWMoE: An Effective Model for Global Weather Forecasting with Mixture-of-ExpertsLihao Gan, Xin Man, Chenghong Zhang, Jie ShaoAAAI 2025 · 10 citations
- Scaling Law for Time Series ForecastingJingzhe Shi, Qinwei Ma, Huan Ma, Lei LiNeurIPS 2024 · 39 citations
- Fixing the Double Penalty in Data-Driven Weather Forecasting Through a Modified Spherical Harmonic Loss FunctionChristopher Subich, Syed Zahid Husain, Leo Separovic, Jing YangICML 2025
- Pre-training under infinite computeKonwoo Kim, Suhas Kotha, Percy Liang, Tatsunori HashimotoICLR 2026 · 25 citations
- Towards Precise Scaling Laws for Video Diffusion TransformersYuanyang Yin, Yaqi Zhao, Mingwu Zheng, Ke Lin et al.CVPR 2025
