Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
Xu Liu, Juncheng Liu, Gerald Woo, Taha Aksu, Yuxuan Liang, Roger Zimmermann, Chenghao Liu, Junnan Li, Silvio Savarese, Caiming Xiong, Doyen Sahoo
摘要
Time series foundation models have demonstrated impressive performance as zeroshot forecasters, i.e., they can tackle a wide variety of downstream forecasting tasks without explicit task-specific training. However, achieving effectively unified training on time series remains an open challenge. Existing approaches introduce some level of model specialization to account for the highly heterogeneous nature of time series data. For instance, MOIRAI pursues unified training by employing multiple input/output projection layers, each tailored to handle time series at a specific frequency. Similarly, TimesFM maintains a frequency embedding dictionary for this purpose. We identify two major drawbacks to this human-imposed frequency-level model specialization: (1) Frequency is not a reliable indicator of the underlying patterns in time series. For example, time series with different frequencies can display similar patterns, while those with the same frequency may exhibit varied patterns. (2) Non-stationarity is an inherent property of real-world time series, leading to varied distributions even within a short context window of a single time series. Frequency-level specialization is too coarse-grained to capture this level of diversity. To address these limitations, this paper introduces MOIRAI-MOE, using a single input/output projection layer while delegating the modeling of diverse time series patterns to the sparse mixture of experts (MoE) within Transformers. With these designs, MOIRAI-MOE reduces reliance on human-defined heuristics and enables automatic token-level specialization. Extensive experiments on 39 datasets demonstrate the superiority of MOIRAI-MOE over existing foundation models in both in-distribution and zero-shot scenarios. Furthermore, this study conducts comprehensive model analyses to explore the inner workings of time series MoE foundation models and provides valuable insights for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- This Time is Different: An Observability Perspective on Time Series Foundation ModelsBen Cohen, Emaad Khwaja, Youssef Doubli, Salahidine Lemaachi 等NeurIPS 2025 · 被引用 68 次
- MIRA: Medical Time Series Foundation Model for Real-World Health DataHao Li, Bowen Deng, Chang Xu, Zhiyuan Feng 等NeurIPS 2025 · 被引用 27 次
- CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic DataShifeng Xie, Vasilii Feofanov, Jianfeng Zhang, Themis Palpanas 等ICLR 2026 · 被引用 15 次
- Attention Mechanism, Max-Affine Partition, and Universal ApproximationHude Liu, Jerry Yao-Chieh Hu, Zhao Song, Han LiuNeurIPS 2025 · 被引用 12 次
- SEMPO: Lightweight Foundation Models for Time Series ForecastingHui He, Kun Yi, Yuanchi Ma, Qi Zhang 等NeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper18
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
相关 Paper
- Multi-Scale Finetuning for Encoder-based Time Series Foundation ModelsZhongzheng Qiao, Chenghao Liu, Yiming Zhang, Ming Jin 等NeurIPS 2025 · 被引用 17 次
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong 等ICML 2024 · 被引用 513 次
- Zeus: Towards Tuning-Free Foundation Model for Time Series AnalysisYisong Fu, Zezhi Shao, Chengqing Yu, Yujie Li 等ICML 2026
- Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of ExpertsXiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li 等ICLR 2025
- In-Context Fine-Tuning for Time-Series Foundation ModelsMatthew Faw, Rajat Sen, Yichen Zhou, Abhimanyu DasICML 2025
