Towards Neural Scaling Laws for Time Series Foundation Models
Qingren Yao, Chao-Han Huck Yang, Renhe Jiang, Yuxuan Liang, Ming Jin, Shirui Pan
Abstract
Scaling laws offer valuable insights into the design of time series foundation models (TSFMs). However, previous research has largely focused on the scaling laws of TSFMs for in-distribution (ID) data, leaving their out-of-distribution (OOD) scaling behavior and the influence of model architectures less explored. In this work, we examine two common TSFM architectures-encoder-only and decoderonly Transformers-and investigate their scaling behavior on both ID and OOD data. These models are trained and evaluated across varying parameter counts, compute budgets, and dataset sizes. Our experiments reveal that the negative loglikelihood of TSFMs exhibits similar scaling behavior in both OOD and ID settings. We further compare the scaling properties across different architectures, incorporating two state-of-the-art TSFMs as case studies, showing that model architecture plays a significant role in scaling. The encoder-only Transformers demonstrate better scalability than the decoder-only Transformers in ID data, while the architectural enhancements in the two advanced TSFMs primarily improve ID performance but reduce OOD scalability. While scaling up TSFMs is expected to drive performance breakthroughs, the lack of a comprehensive understanding of TSFM scaling laws has hindered the development of a robust framework to guide model scaling. We fill this gap in this work by synthesizing our findings and providing practical guidelines for designing and scaling larger TSFMs with enhanced model capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- This Time is Different: An Observability Perspective on Time Series Foundation ModelsBen Cohen, Emaad Khwaja, Youssef Doubli, Salahidine Lemaachi et al.NeurIPS 2025 · 68 citations
- Aurora: Towards Universal Generative Multimodal Time Series ForecastingXingjian Wu, Jianxin Jin, Wanghui Qiu, Peng Chen et al.ICLR 2026 · 33 citations
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language ModelsTong Guan, Zijie Meng, Dianqi Li, Shiyu Wang et al.ICLR 2026 · 29 citations
- Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learningYuanzhao Zhang, William GilpinICLR 2026 · 16 citations
- CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic DataShifeng Xie, Vasilii Feofanov, Jianfeng Zhang, Themis Palpanas et al.ICLR 2026 · 15 citations
Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 601 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
Related papers
- Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of ExpertsXiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li et al.ICLR 2025
- Scaling Law for Time Series ForecastingJingzhe Shi, Qinwei Ma, Huan Ma, Lei LiNeurIPS 2024 · 39 citations
- Scaling View Synthesis TransformersEvan Kim, Hyunwoo Ryu, Thomas W. Mitchel, Vincent SitzmannCVPR 2026 · 6 citations
- FlowState: Sampling-Rate‑Equivariant Time‑Series ForecastingLars Graf, Thomas Ortner, Stanisław Woźniak, Angeliki PantaziICML 2026
- Revealing Scaling Paradox in Large-scale Time Series Models: Implications for More Efficient and Accurate ForecastingXin Qiu, Junlong Tong, Yirong Sun, Yunpu Ma et al.ICML 2026
