Are Self-Attentions Effective for Time Series Forecasting?
Dongbin Kim, Jinseong Park, Jaewook Lee, Hoki Kim
摘要
Time series forecasting is crucial for applications across multiple domains and various scenarios. Although Transformer models have dramatically advanced the landscape of forecasting, their effectiveness remains debated. Recent findings have indicated that simpler linear models might outperform complex Transformer-based approaches, highlighting the potential for more streamlined architectures. In this paper, we shift the focus from evaluating the overall Transformer architecture to specifically examining the effectiveness of self-attention for time series forecasting. To this end, we introduce a new architecture, Cross-Attention-only Time Series transformer (CATS), that rethinks the traditional Transformer framework by eliminating self-attention and leveraging cross-attention mechanisms instead. By establishing future horizon-dependent parameters as queries and enhanced parameter sharing, our model not only improves long-term forecasting accuracy but also reduces the number of parameters and memory usage. Extensive experiment across various datasets demonstrates that our model achieves superior performance with the lowest mean squared error and uses fewer parameters compared to existing models. The implementation of our model is available at: https://github.com/dongbeank/CATS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Towards Robust Real-World Multivariate Time Series Forecasting: A Unified Framework for Dependency, Asynchrony, and MissingnessJinkwan Jang, Hyungjin Park, Jinmyeong Choi, Taesup KimICLR 2026 · 被引用 2 次
- EMAformer: Enhancing Transformer Through Embedding Armor for Time Series ForecastingZhiwei Zhang, Xinyi Du, Xuanchi Guo, Weihao Wang 等AAAI 2026 · 被引用 1 次
- TimePerceiver: An Encoder-Decoder Framework for Generalized Time-Series ForecastingJaebin Lee, Hankook LeeNeurIPS 2025 · 被引用 1 次
- M2FMoE: Multi-Resolution Multi-View Frequency Mixture-of-Experts for Extreme-Adaptive Time Series ForecastingYaohui Huang, Runmin Zou, Yun Wang, Laeeq Aslam 等AAAI 2026
- TEDM: Time Series Forecasting with Elucidated Diffusion ModelsEdgardo Solano-Carrillo, Sreerag Vadakkemeppully Naveenachandran, Julia NieblingICLR 2026
它引用的顶会 Paper20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 被引用 3,619 次
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang 等ICML 2022 · 被引用 2,912 次
相关 Paper
- Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?Xinyu Li, Kexi Chen, Jiajie Shen, Ying Zheng 等ACL 2026
- Unlocking the Power of Patch: Patch-Based MLP for Long-Term Time Series ForecastingPeiwang Tang, Weitai ZhangAAAI 2025 · 被引用 42 次
- SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise AttentionRomain Ilbert, Ambroise Odonnat, Vasilii Feofanov, Aladin Virmaux 等ICML 2024 · 被引用 62 次
- A Lightweight Sparse Interaction Network for Time Series ForecastingXu Zhang, Qitong Wang, Peng Wang, Wei WangAAAI 2025 · 被引用 1 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
