Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts
Xiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li, Zhou Ye, Qingsong Wen, Ming Jin
Abstract
Deep learning for time series forecasting has seen significant advancements over the past decades. However, despite the success of large-scale pre-training in language and vision domains, pre-trained time series models remain limited in scale and operate at a high cost, hindering the development of larger capable forecasting models in real-world applications. In response, we introduce Time-MoE, a scalable and unified architecture designed to pre-train larger, more capable forecasting foundation models while reducing inference costs. By leveraging a sparse mixture-of-experts (MoE) design, Time-MoE enhances computational efficiency by activating only a subset of networks for each prediction, reducing computational load while maintaining high model capacity. This allows Time-MoE to scale effectively without a corresponding increase in inference costs. Time-MoE comprises a family of decoder-only transformer models that operate in an auto-regressive manner and support flexible forecasting horizons with varying input context lengths. We pre-trained these models on our newly introduced large-scale data Time-300B, which spans over 9 domains and encompassing over 300 billion time points. For the first time, we scaled a time series foundation model up to 2.4 billion parameters, achieving significantly improved forecasting precision. Our results validate the applicability of scaling laws for training tokens and model size in the context of time series forecasting. Compared to dense models with the same number of activated parameters or equivalent computation budgets, our models consistently outperform them by large margin. These advancements position Time-MoE as a state-of-the-art solution for tackling real-world time series forecasting challenges with superior capability, efficiency, and flexibility. Code is available at https://github.com/Time-MoE/Time-MoE
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5da77dbb-16d8-4590-906a-63d36440b1e6Cited by top-tier papers81
- This Time is Different: An Observability Perspective on Time Series Foundation ModelsBen Cohen, Emaad Khwaja, Youssef Doubli, Salahidine Lemaachi et al.NeurIPS 2025 · 68 citations
- TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot ForecasterKanghui Ning, Zijie Pan, Yu Liu, Yushan Jiang et al.NeurIPS 2025 · 53 citations
- OLinear: A Linear Model for Time Series Forecasting in Orthogonally Transformed DomainWenzhen Yue, Yong Liu, Hao Wang, Haoxuan Li et al.NeurIPS 2025 · 42 citations
- Aurora: Towards Universal Generative Multimodal Time Series ForecastingXingjian Wu, Jianxin Jin, Wanghui Qiu, Peng Chen et al.ICLR 2026 · 33 citations
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language ModelsTong Guan, Zijie Meng, Dianqi Li, Shiyu Wang et al.ICLR 2026 · 29 citations
Builds on28
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong et al.ICML 2024 · 513 citations
- SEMPO: Lightweight Foundation Models for Time Series ForecastingHui He, Kun Yi, Yuanchi Ma, Qi Zhang et al.NeurIPS 2025 · 12 citations
- Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of ExpertsXu Liu, Juncheng Liu, Gerald Woo, Taha Aksu et al.ICML 2025
- Sundial: A Family of Highly Capable Time Series Foundation ModelsYong Liu, Guo Qin, Zhiyuan Shi, Zhi Chen et al.ICML 2025
- Timer: Generative Pre-trained Transformers Are Large Time Series ModelsYong Liu, Haoran Zhang, Chenyu Li, Xiangdong Huang et al.ICML 2024 · 188 citations
