Scaling Law for Time Series Forecasting
Jingzhe Shi, Qinwei Ma, Huan Ma, Lei Li
Abstract
Scaling law that rewards large datasets, complex models and enhanced data granularity has been observed in various fields of deep learning. Yet, studies on time series forecasting have cast doubt on scaling behaviors of deep learning methods for time series forecasting: while more training data improves performance, more capable models do not always outperform less capable models, and longer input horizons may hurt performance for some models. We propose a theory for scaling law for time series forecasting that can explain these seemingly abnormal behaviors. We take into account the impact of dataset size and model complexity, as well as time series data granularity, particularly focusing on the look-back horizon, an aspect that has been unexplored in previous theories. Furthermore, we empirically evaluate various models using a diverse set of time series forecasting datasets, which (1) verifies the validity of scaling law on dataset size and model complexity within the realm of time series forecasting, and (2) validates our theoretical framework, particularly regarding the influence of look back horizon. We hope our findings may inspire new models targeting time series forecasting datasets of limited size, as well as large foundational datasets and models for time series forecasting in future work. Code for our experiments has been made public at https://github.com/JingzheShi/ScalingLawForTimeSeriesForecasting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- TAB: Unified Benchmarking of Time Series Anomaly Detection MethodsXiangfei Qiu, Zhe Li, Wanghui Qiu, Shiyan Hu et al.VLDB 2025 · 57 citations
- Aurora: Towards Universal Generative Multimodal Time Series ForecastingXingjian Wu, Jianxin Jin, Wanghui Qiu, Peng Chen et al.ICLR 2026 · 33 citations
- Intrinsic Entropy of Context Length Scaling in LLMsJingzhe Shi, Qinwei Ma, Hongyi Liu, Hang Zhao et al.ICLR 2026 · 17 citations
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning ModelsYuhui Wang, Changjiang Li, Guangke Chen, Jiacheng Liang et al.ICLR 2026 · 13 citations
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou et al.CVPR 2026 · 12 citations
Builds on18
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
Related papers
- Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of ExpertsXiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li et al.ICLR 2025
- Scaling Laws of Global Weather ModelsYuejiang Yu, Langwen Huang, Alexandru Calotoiu, Torsten HoeflerICML 2026 · 4 citations
- TimeRecipe: A Time-Series Forecasting Recipe via Benchmarking Module Level EffectivenessZhiyuan Zhao, Juntong Ni, Shangqing Xu, Haoxin Liu et al.ICLR 2026 · 7 citations
- Scaleformer: Iterative Multi-scale Refining Transformers for Time Series ForecastingMohammad Amin Shabani, Amir H. Abdi, Lili Meng, Tristan SylvainICLR 2023 · 36 citations
- Towards Neural Scaling Laws for Time Series Foundation ModelsQingren Yao, Chao-Han Huck Yang, Renhe Jiang, Yuxuan Liang et al.ICLR 2025
