Pyraformer: Low-Complexity Pyramidal Attention for Long-Range Time Series Modeling and Forecasting
Shizhan Liu, Hang Yu, Cong Liao, Jianguo Li, Weiyao Lin, Alex X. Liu, Schahram Dustdar
Abstract
Accurate prediction of the future given the past based on time series data is of paramount importance, since it opens the door for decision making and risk management ahead of time. In practice, the challenge is to build a flexible but parsimonious model that can capture a wide range of temporal dependencies. In this paper, we propose Pyraformer by exploring the multi-resolution representation of the time series. Specifically, we introduce the pyramidal attention module (PAM) in which the inter-scale tree structure summarizes features at different resolutions and the intra-scale neighboring connections model the temporal dependencies of different ranges. Under mild conditions, the maximum length of the signal traversing path in Pyraformer is a constant (i.e., O(1)) with regard to the sequence length L, while its time and space complexity scale linearly with L. Extensive experimental results show that Pyraformer typically achieves the highest prediction accuracy in both single-step and long-range multi-step forecasting tasks with the least amount of time and memory consumption, especially when the sequence is long 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac2b7b2c-024d-4ee2-a54f-42752fddebebCited by top-tier papers194
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun et al.NeurIPS 2023 · 1,178 citations
- Non-stationary Transformers: Exploring the Stationarity in Time Series ForecastingYong Liu, Haixu Wu, Jianmin Wang, Mingsheng LongNeurIPS 2022 · 1,080 citations
- SCINet: Time Series Modeling and Forecasting with Sample Convolution and InteractionMinhao Liu, Ailing Zeng, Muxi Chen, Zhijian Xu et al.NeurIPS 2022 · 934 citations
Builds on5
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformer Hawkes ProcessSimiao Zuo, Haoming Jiang, Zichong Li, Tuo Zhao et al.ICML 2020 · 382 citations
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek et al.EMNLP 2020 · 268 citations
Related papers
- Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series ForecastingPeng Chen, Yingying Zhang, Yunyao Cheng, Yang Shu et al.ICLR 2024 · 197 citations
- Autohformer: Efficient Hierarchical Autoregressive Transformer for Time Series PredictionQianru Zhang, Honggang Wen, Ming Li, Dong Huang et al.ICDE 2026
- HMformer: Unleashing Transformer's Potential for Time Series Forecasting via Hierarchical Multi-Scale ModelingRenjun Huang, Han Xiao, Bingqing Li, Baili Zhang et al.AAAI 2026
- Peri-midFormer: Periodic Pyramid Transformer for Time Series AnalysisQiang Wu, Gechang Yao, Zhixi Feng, Shuyuan YangNeurIPS 2024 · 21 citations
- Sparse-Scale Transformer with Bidirectional Awareness for Time Series ForecastingYing Liu, Bo Liu, Sheng Huang, Gang Luo et al.AAAI 2026
