HDT: Hierarchical Discrete Transformer for Multivariate Time Series Forecasting
Shibo Feng, Peilin Zhao, Liu Liu, Pengcheng Wu, Zhiqi Shen
Abstract
Generative models have gained significant attention in multivariate time series forecasting (MTS), particularly due to their ability to generate high-fidelity samples. Forecasting the probability distribution of multivariate time series is a challenging yet practical task. Although some recent attempts have been made to handle this task, two major challenges persist: 1) some existing generative methods underperform in high-dimensional multivariate time series forecasting, which is hard to scale to higher dimensions; 2) The inherent high-dimensional multivariate attributes constrain the forecasting lengths of existing generative models. In this paper, we point out that discrete token representations can model high-dimensional MTS with faster inference time, and forecast the target with the long-term trends of itself can extend the forecasting length with high accuracy. Motivated by this, we propose a vector quantized framework called Hierarchical Discrete Transformer (HDT) that models time series into discrete token representations with l2 normalization enhanced vector quantized strategy, in which we transform the MTS forecasting into discrete tokens generation. To address the limitations of generative models in long-term forecasting, we propose a hierarchical discrete Transformer. This model captures the discrete long-term trend of the target at the low level and leverages this trend as a condition to generate the discrete representation of the target at the high level that introduces the features of target itself for extending the forecasting length in high-dimensional MTS. Extensive experiments on five popular MTS datasets verify the effectiveness of our proposed method. The source code will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7339b504-2570-4d2c-8aa3-4084b23a5165Cited by top-tier papers5
- Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and CompressibilityAnnan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang et al.ICLR 2026 · 9 citations
- TS-LIF: A Temporal Segment Spiking Neuron Network for Time Series ForecastingShibo Feng, Wanjin Feng, Xingyu Gao, Peilin Zhao et al.ICLR 2025
- Hi-Time: Hierarchical Latent Prediction for Multivariate Time Series ClassificationKun Zeng, Wu Binquan, Qianli MaICML 2026
- ReCast: Reliability-aware Codebook-assisted Lightweight Time Series ForecastingXiang Ma, Taihua Chen, Pengcheng Wang, Xuemei Li et al.AAAI 2026
- CrystalDiT: Simple Diffusion Transformers for Crystal GenerationXiaohan Yi, Guikun Xu, Zhong Zhang, Liu Liu et al.AAAI 2026
Builds on24
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- CSDI: Conditional Score-based Diffusion Models for Probabilistic Time Series ImputationYusuke Tashiro, Jiaming Song, Yang Song, Stefano ErmonNeurIPS 2021 · 1,245 citations
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun et al.NeurIPS 2023 · 1,178 citations
Related papers
- Considering Nonstationary within Multivariate Time Series with Variational Hierarchical Transformer for ForecastingMuyao Wang, Wenchao Chen, Bo ChenAAAI 2024 · 13 citations
- SDformer: Similarity-driven Discrete Transformer For Time Series GenerationZhicheng Chen, Shibo Feng, Zhong Zhang, Xi Xiao et al.NeurIPS 2024 · 28 citations
- VQ-TR: Vector Quantized Attention for Time Series ForecastingKashif Rasul, Andrew Bennett, Pablo Vicente, Umang Gupta et al.ICLR 2024 · 7 citations
- Latent Diffusion Transformer for Probabilistic Time Series ForecastingShibo Feng, Chunyan Miao, Zhong Zhang, Peilin ZhaoAAAI 2024 · 60 citations
- Probabilistic Transformer For Time Series AnalysisBinh Tang, David S. MattesonNeurIPS 2021 · 150 citations
