Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts
Xiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li, Zhou Ye, Qingsong Wen, Ming Jin
摘要
Deep learning for time series forecasting has seen significant advancements over the past decades. However, despite the success of large-scale pre-training in language and vision domains, pre-trained time series models remain limited in scale and operate at a high cost, hindering the development of larger capable forecasting models in real-world applications. In response, we introduce Time-MoE, a scalable and unified architecture designed to pre-train larger, more capable forecasting foundation models while reducing inference costs. By leveraging a sparse mixture-of-experts (MoE) design, Time-MoE enhances computational efficiency by activating only a subset of networks for each prediction, reducing computational load while maintaining high model capacity. This allows Time-MoE to scale effectively without a corresponding increase in inference costs. Time-MoE comprises a family of decoder-only transformer models that operate in an auto-regressive manner and support flexible forecasting horizons with varying input context lengths. We pre-trained these models on our newly introduced large-scale data Time-300B, which spans over 9 domains and encompassing over 300 billion time points. For the first time, we scaled a time series foundation model up to 2.4 billion parameters, achieving significantly improved forecasting precision. Our results validate the applicability of scaling laws for training tokens and model size in the context of time series forecasting. Compared to dense models with the same number of activated parameters or equivalent computation budgets, our models consistently outperform them by large margin. These advancements position Time-MoE as a state-of-the-art solution for tackling real-world time series forecasting challenges with superior capability, efficiency, and flexibility. Code is available at https://github.com/Time-MoE/Time-MoE
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper81
- This Time is Different: An Observability Perspective on Time Series Foundation ModelsBen Cohen, Emaad Khwaja, Youssef Doubli, Salahidine Lemaachi 等NeurIPS 2025 · 被引用 68 次
- TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot ForecasterKanghui Ning, Zijie Pan, Yu Liu, Yushan Jiang 等NeurIPS 2025 · 被引用 53 次
- OLinear: A Linear Model for Time Series Forecasting in Orthogonally Transformed DomainWenzhen Yue, Yong Liu, Hao Wang, Haoxuan Li 等NeurIPS 2025 · 被引用 42 次
- Aurora: Towards Universal Generative Multimodal Time Series ForecastingXingjian Wu, Jianxin Jin, Wanghui Qiu, Peng Chen 等ICLR 2026 · 被引用 33 次
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language ModelsTong Guan, Zijie Meng, Dianqi Li, Shiyu Wang 等ICLR 2026 · 被引用 29 次
它引用的顶会 Paper28
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 被引用 3,619 次
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang 等ICML 2022 · 被引用 2,912 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong 等ICML 2024 · 被引用 513 次
- SEMPO: Lightweight Foundation Models for Time Series ForecastingHui He, Kun Yi, Yuanchi Ma, Qi Zhang 等NeurIPS 2025 · 被引用 12 次
- Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of ExpertsXu Liu, Juncheng Liu, Gerald Woo, Taha Aksu 等ICML 2025
- Sundial: A Family of Highly Capable Time Series Foundation ModelsYong Liu, Guo Qin, Zhiyuan Shi, Zhi Chen 等ICML 2025
- Timer: Generative Pre-trained Transformers Are Large Time Series ModelsYong Liu, Haoran Zhang, Chenyu Li, Xiangdong Huang 等ICML 2024 · 被引用 188 次
