A Closer Look at Transformers for Time Series Forecasting: Understanding Why They Work and Where They Struggle
Yu Chen, Nathalia Céspedes, Payam M. Barnaghi
摘要
Time-series forecasting is crucial across various domains, including finance, healthcare, and energy. Transformer models, originally developed for natural language processing, have demonstrated significant potential in addressing challenges associated with time-series data. These models utilize different tokenization strategies, point-wise, patch-wise, and variate-wise, to represent time-series data, each resulting in different scope of attention maps. Despite the emergence of sophisticated architectures, simpler transformers consistently outperform their more complex counterparts in widely used benchmarks. This study examines why point-wise transformers are generally less effective, why intra-and inter-variate attention mechanisms yield similar outcomes, and which architectural components drive the success of simpler models. By analyzing mutual information and evaluating models on synthetic datasets, we demonstrate that intravariate dependencies are the primary contributors to prediction performance on benchmarks, while inter-variate dependencies have a minor impact. Additionally, techniques such as Z-score normalization and skip connections are also crucial. However, these results are largely influenced by the self-dependent and stationary nature of benchmark datasets. By validating our findings on real-world healthcare data, we provide insights for designing more effective transformers for practical applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DropoutTS: Sample-Adaptive Dropout for Robust Time Series ForecastingSiru Zhong, Yiqiu Liu, Zhiqing Cui, Zezhi Shao 等ICML 2026
- Byte Pair Encoding for Efficient Time Series ForecastingLeon Götz, Marcel Kollovieh, Stephan Günnemann, Leo SchwinnICML 2026
- Revealing Scaling Paradox in Large-scale Time Series Models: Implications for More Efficient and Accurate ForecastingXin Qiu, Junlong Tong, Yirong Sun, Yunpu Ma 等ICML 2026
- CombinationTS: A Modular Framework for Understanding Time-Series Forecasting ModelsXiaorui Wang, Fanda Fan, Chenxi Wang, Yuxuan Yang 等ICML 2026
它引用的顶会 Paper6
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
- Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution ShiftTaesung Kim, Jinhee Kim, Yunwon Tae, Cheonbok Park 等ICLR 2022 · 被引用 1,020 次
- Pyraformer: Low-Complexity Pyramidal Attention for Long-Range Time Series Modeling and ForecastingShizhan Liu, Hang Yu, Cong Liao, Jianguo Li 等ICLR 2022 · 被引用 975 次
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 被引用 536 次
相关 Paper
- SimpleTM: A Simple Baseline for Multivariate Time Series ForecastingHui Chen, Viet Luong, Lopamudra Mukherjee, Vikas SinghICLR 2025
- InParformer: Evolutionary Decomposition Transformers with Interactive Parallel Attention for Long-Term Time Series ForecastingHaizhou Cao, Zhenhao Huang, Tiechui Yao, Jue Wang 等AAAI 2023 · 被引用 29 次
- Sequence Complementor: Complementing Transformers for Time Series Forecasting with Learnable SequencesXiwen Chen, Peijie Qiu, Wenhui Zhu, Huayu Li 等AAAI 2025 · 被引用 4 次
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 被引用 3,619 次
- Sparse-Scale Transformer with Bidirectional Awareness for Time Series ForecastingYing Liu, Bo Liu, Sheng Huang, Gang Luo 等AAAI 2026
