Understanding Transformers for Time Series Forecasting: A Case Study on MOIRAI
Yu-Hsuan Wu, Yihan He, Yuan Cao, Jianqing Fan, Han Liu
摘要
We give a comprehensive theoretical analysis of transformers as time series prediction models, with a focus on MOIRAI (Woo et al., 2024) . We study its approximation and generalization capabilities. First, we demonstrate that there exist transformers that fit an autoregressive model on input univariate time series via gradient descent. We then analyze MOIRAI, one of the state-of-the-art multivariate time series prediction models capable of modeling arbitrary number of covariates. We prove that MOIRAI is capable of automatically fitting autoregressive models with an arbitrary number of covariates, offering insights into its design and empirical success. For generalization, we establish learning bounds for pretraining when the data satisfies Dobrushin's condition. Experiments support our theoretical findings, highlighting the efficacy of using transformers for time series forecasting. * equal contribution PROBLEM SETUP This section presents backgrounds and formula definitions of the transformer model, and then introduce the auto-regressive model. Transformers. We consider a sequence of N input vectors h Given any H ∈ R D×N , we define the attention layer as follows. Definition 2.1 (Attention layer). A self-attention layer with M heads is denoted as Attn . The self-attention layer processes any given input sequence H ∈ R D×N as Attn † θ 0 (H) := H + 1 N M m=1 (VmH) × σ (QmH) ⊤ (KmH) , where σ(t) := ReLU(t)/N is the ReLU function normalized by N . Next, we introduce the any-variate attention, where Woo et al. (2024) uses it to replace the standard attention in transformers. The any-variate attention introduces two learnable variables: Attention
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu 等ICLR 2024 · 被引用 915 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
相关 Paper
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong 等ICML 2024 · 被引用 513 次
- SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise AttentionRomain Ilbert, Ambroise Odonnat, Vasilii Feofanov, Aladin Virmaux 等ICML 2024 · 被引用 62 次
- Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of ExpertsXu Liu, Juncheng Liu, Gerald Woo, Taha Aksu 等ICML 2025
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty 等KDD 2021 · 被引用 66 次
- Timer-XL: Long-Context Transformers for Unified Time Series ForecastingYong Liu, Guo Qin, Xiangdong Huang, Jianmin Wang 等ICLR 2025
