Understanding Transformers for Time Series Forecasting: A Case Study on MOIRAI
Yu-Hsuan Wu, Yihan He, Yuan Cao, Jianqing Fan, Han Liu
Abstract
We give a comprehensive theoretical analysis of transformers as time series prediction models, with a focus on MOIRAI (Woo et al., 2024) . We study its approximation and generalization capabilities. First, we demonstrate that there exist transformers that fit an autoregressive model on input univariate time series via gradient descent. We then analyze MOIRAI, one of the state-of-the-art multivariate time series prediction models capable of modeling arbitrary number of covariates. We prove that MOIRAI is capable of automatically fitting autoregressive models with an arbitrary number of covariates, offering insights into its design and empirical success. For generalization, we establish learning bounds for pretraining when the data satisfies Dobrushin's condition. Experiments support our theoretical findings, highlighting the efficacy of using transformers for time series forecasting. * equal contribution PROBLEM SETUP This section presents backgrounds and formula definitions of the transformer model, and then introduce the auto-regressive model. Transformers. We consider a sequence of N input vectors h Given any H ∈ R D×N , we define the attention layer as follows. Definition 2.1 (Attention layer). A self-attention layer with M heads is denoted as Attn . The self-attention layer processes any given input sequence H ∈ R D×N as Attn † θ 0 (H) := H + 1 N M m=1 (VmH) × σ (QmH) ⊤ (KmH) , where σ(t) := ReLU(t)/N is the ReLU function normalized by N . Next, we introduce the any-variate attention, where Woo et al. (2024) uses it to replace the standard attention in transformers. The any-variate attention introduces two learnable variables: Attention
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbc49015-36d1-41c5-aa72-e12e8a5ee0b1Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
Related papers
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong et al.ICML 2024 · 513 citations
- SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise AttentionRomain Ilbert, Ambroise Odonnat, Vasilii Feofanov, Aladin Virmaux et al.ICML 2024 · 62 citations
- Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of ExpertsXu Liu, Juncheng Liu, Gerald Woo, Taha Aksu et al.ICML 2025
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty et al.KDD 2021 · 66 citations
- Timer-XL: Long-Context Transformers for Unified Time Series ForecastingYong Liu, Guo Qin, Xiangdong Huang, Jianmin Wang et al.ICLR 2025
