Autohformer: Efficient Hierarchical Autoregressive Transformer for Time Series Prediction
Qianru Zhang, Honggang Wen, Ming Li, Dong Huang, Siu-Ming Yiu, Christian S. Jensen, Pietro Liò
Abstract
Time series forecasting requires architectures that simultaneously achieve three competing objectives: (1) strict temporal causality for reliable predictions, (2) subquadratic complexity for practical scalability, and (3) multiscale pattern recognition for accurate long-horizon forecasting. We introduce AutoHFormer, a hierarchical autoregressive transformer that addresses these challenges through three key innovations: 1) Hierarchical Temporal Modeling: Our architecture decomposes predictions into segment-level blocks processed in parallel, followed by intra-segment sequential refinement. This dual-scale approach maintains temporal coherence while enabling efficient computation. 2) Dynamic Windowed Attention: The attention mechanism employs learnable causal windows with exponential decay, reducing complexity while preserving precise temporal relationships. This design avoids both the anti-causal violations of standard transformers and the sequential bottlenecks of RNN hybrids. 3) Adaptive Temporal Encoding: a novel position encoding system is adopted to capture time patterns at multiple scales. It combines fixed oscillating patterns for short-term variations with learnable decay rates for long-term trends. Comprehensive experiments demonstrate that AutoHFormer 10.76× faster training and 6.06× memory reduction compared to PatchTST on PEMS08, while maintaining consistent accuracy across 96-720 step horizons in most of cases. These breakthroughs establish new benchmarks for efficient and precise time series modeling. Implementations of our method and all baselines in hierarchical autoregressive mechanism are available at https://github.com/CoderPowerBeyond/AutoHFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1d4ceed-a85f-46f4-8c68-06fe147533faCited by top-tier papers1
Ask how each one uses itBuilds on15
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
Related papers
- HMformer: Unleashing Transformer's Potential for Time Series Forecasting via Hierarchical Multi-Scale ModelingRenjun Huang, Han Xiao, Bingqing Li, Baili Zhang et al.AAAI 2026
- Sparse-Scale Transformer with Bidirectional Awareness for Time Series ForecastingYing Liu, Bo Liu, Sheng Huang, Gang Luo et al.AAAI 2026
- Scaleformer: Iterative Multi-scale Refining Transformers for Time Series ForecastingMohammad Amin Shabani, Amir H. Abdi, Lili Meng, Tristan SylvainICLR 2023 · 36 citations
- Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series ForecastingPeng Chen, Yingying Zhang, Yunyao Cheng, Yang Shu et al.ICLR 2024 · 197 citations
- Ada-MSHyper: Adaptive Multi-Scale Hypergraph Transformer for Time Series ForecastingZongjiang Shang, Ling Chen, Binqing Wu, Dongliang CuiNeurIPS 2024 · 49 citations
