Scaling Laws of Global Weather Models
Yuejiang Yu, Langwen Huang, Alexandru Calotoiu, Torsten Hoefler
摘要
Data-driven models are revolutionizing weather forecasting. To optimize training efficiency and model performance, this paper analyzes empirical scaling laws within this domain. We investigate the relationship between model performance (validation loss) and three key factors: model size (), dataset size (), and compute budget (). Across a range of models, we find that Aurora exhibits the strongest data-scaling behavior: increasing the training dataset by 10× reduces validation loss by up to 3.2×. GraphCast demonstrates the highest parameter efficiency, yet suffers from limited hardware utilization. Our compute-optimal analysis indicates that, under fixed compute budgets, allocating resources to more total training data yields greater performance gains than increasing model size. Furthermore, we analyze model shape and uncover scaling behaviors that differ fundamentally from those observed in language models: weather forecasting models consistently favor increased width over depth. These findings suggest that future weather models should prioritize wider architectures and larger effective training datasets to maximize predictive performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Spherical Fourier Neural Operators: Learning Stable Dynamics on the SphereBoris Bonev, Thorsten Kurth, Christian Hundt, Jaideep Pathak 等ICML 2023 · 被引用 280 次
- Scaling transformer neural networks for skillful and reliable medium-range weather forecastingTung Nguyen, Rohan Shah, Hritik Bansal, Troy Arcomano 等NeurIPS 2024 · 被引用 165 次
- Scaling Laws for a Multi-Agent Reinforcement Learning ModelOren Neumann, Claudius GrosICLR 2023 · 被引用 3 次
相关 Paper
- EWMoE: An Effective Model for Global Weather Forecasting with Mixture-of-ExpertsLihao Gan, Xin Man, Chenghong Zhang, Jie ShaoAAAI 2025 · 被引用 10 次
- Scaling Law for Time Series ForecastingJingzhe Shi, Qinwei Ma, Huan Ma, Lei LiNeurIPS 2024 · 被引用 39 次
- Fixing the Double Penalty in Data-Driven Weather Forecasting Through a Modified Spherical Harmonic Loss FunctionChristopher Subich, Syed Zahid Husain, Leo Separovic, Jing YangICML 2025
- Pre-training under infinite computeKonwoo Kim, Suhas Kotha, Percy Liang, Tatsunori HashimotoICLR 2026 · 被引用 25 次
- Towards Precise Scaling Laws for Video Diffusion TransformersYuanyang Yin, Yaqi Zhao, Mingwu Zheng, Ke Lin 等CVPR 2025
