Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?
Xinyu Li, Kexi Chen, Jiajie Shen, Ying Zheng, Hong Lu, Jin Zhao, Xin Wang
Abstract
In multivariate long-term time series forecasting, the success of self-attention is commonly attributed to the attention matrix that encodes token interactions. In this paper, we provide evidence that challenges this view. Through extensive experiments on three classic and three latest Transformer models, we find that dotproduct attention can be replaced by elementwise operations without token interaction, such as the addition and Hadamard product, while maintaining or even improving accuracy. This motivates our central hypothesis: the effectiveness of self-attention in this task arises not from the dynamic attention matrix, but from the multi-branch feature extraction enabled by the parallel Query, Key, and Value projections and their fusion. To validate this hypothesis, we construct a minimalist multi-branch MLP that isolates the 'multi-branch mapping with element-wise operation' structure from the Transformer and show that it achieves competitive performance. Our findings indicate that the source of performance in self-attention is often misinterpreted, as its actual advantage stems from the architectural principle of multi-branch mapping and fusion, rather than the attention matrix. Source code is available at: https: //github.com/lxy-PhD2022/Attention
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc519802-3bd2-4325-8440-6a8a541433ecBuilds on23
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
Related papers
- Are Self-Attentions Effective for Time Series Forecasting?Dongbin Kim, Jinseong Park, Jaewook Lee, Hoki KimNeurIPS 2024 · 48 citations
- Synthesizer: Rethinking Self-Attention for Transformer ModelsYi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan et al.ICML 2021 · 399 citations
- Unlocking the Power of Patch: Patch-Based MLP for Long-Term Time Series ForecastingPeiwang Tang, Weitai ZhangAAAI 2025 · 42 citations
- Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive ForecastingJiecheng Lu, Shihao YangICML 2025
- Sequence Complementor: Complementing Transformers for Time Series Forecasting with Learnable SequencesXiwen Chen, Peijie Qiu, Wenhui Zhu, Huayu Li et al.AAAI 2025 · 4 citations
