SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention
Romain Ilbert, Ambroise Odonnat, Vasilii Feofanov, Aladin Virmaux, Giuseppe Paolo, Themis Palpanas, Ievgen Redko
Abstract
Transformer-based architectures achieved breakthrough performance in natural language processing and computer vision, yet they remain inferior to simpler linear baselines in multivariate long-term forecasting. To better understand this phenomenon, we start by studying a toy linear forecasting problem for which we show that transformers are incapable of converging to their true solution despite their high expressive power. We further identify the attention of transformers as being responsible for this low generalization capacity. Building upon this insight, we propose a shallow lightweight transformer model that successfully escapes bad local minima when optimized with sharpness-aware optimization. We empirically demonstrate that this result extends to all commonly used real-world multivariate time series datasets. In particular, SAMformer surpasses current state-of-the-art methods and is on par with the biggest foundation model MOIRAI while having significantly fewer parameters. The code is available at https://github.com/romilbert/samformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95e65119-a33d-4e66-bea9-3fa586030fd2Cited by top-tier papers16
- This Time is Different: An Observability Perspective on Time Series Foundation ModelsBen Cohen, Emaad Khwaja, Youssef Doubli, Salahidine Lemaachi et al.NeurIPS 2025 · 68 citations
- EEO-TFV: Escape-Explore Optimizer for Web-Scale Time-Series Forecasting and Vision AnalysisHua Wang, Jinghao Lu, Fan ZhangWWW 2026 · 6 citations
- Not All Data are Good Labels: On the Self-supervised Labeling for Time Series ForecastingYuxuan Yang, Dalin Zhang, Yuxuan Liang, Hua Lu et al.NeurIPS 2025 · 5 citations
- Synthetic Series-Symbol Data Generation for Time Series Foundation ModelsWenxuan Wang, Kai Wu, Yujian Betterest Li, Dan Wang et al.NeurIPS 2025 · 1 citation
- EMAformer: Enhancing Transformer Through Embedding Armor for Time Series ForecastingZhiwei Zhang, Xinyi Du, Xuanchi Guo, Weihao Wang et al.AAAI 2026 · 1 citation
Builds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
Related papers
- Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive ForecastingJiecheng Lu, Shihao YangICML 2025
- Understanding Transformers for Time Series Forecasting: A Case Study on MOIRAIYu-Hsuan Wu, Yihan He, Yuan Cao, Jianqing Fan et al.ICLR 2026
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- Sparse-Scale Transformer with Bidirectional Awareness for Time Series ForecastingYing Liu, Bo Liu, Sheng Huang, Gang Luo et al.AAAI 2026
