Scaling Up Multivariate Time Series Pre-Training with Decoupled Spatial-Temporal Representations
Rui Zha, Le Zhang, Shuangli Li, Jingbo Zhou, Tong Xu, Hui Xiong, Enhong Chen
Abstract
Data scale has been acknowledged as a crucial factor for enhancing the generalization and effectiveness of pre-training models. While existing methods of multivariate time series pre-training are primarily limited to a single specific dataset, scaling to a larger scenario that includes multiple diverse datasets (e.g., multi-region data) remains a substantial challenge. In this paper, we present a novel Decoupled Spatial-Temporal Representation Learning (DeSTR) framework to serve as the backbone network for investigating the data scaling capability of multivariate time series pre-training architectures. Specifically, DeSTR utilizes two separate encoders to capture both the temporal dynamics within each time series and the spatial correlations among multiple variables. The obtained representations of distinct modalities are then fed into a Spatial-Guided Temporal Transformer to equip the temporal features with spatial discriminative information. Moreover, we employ masked autoencoding as the foundational pre-training framework and introduce spacetime-agnostic augmentation to improve robustness and facilitate implicit spatiotemporal modeling. Finally, we successfully pre-train a unified time series representation learning framework on real-world datasets from three different cities. Extensive experiments are carried out on various downstream tasks to validate the performance of DeSTR, compared with three categories of state-of-the-art baselines: deep sequential models, spatial-temporal graph neural networks, and time series representation learning methods. The results clearly demonstrate the advantages of scaling multivariate time series pre-training to multiple datasets, highlighting the effectiveness of DeSTR as a general spatiotemporal learner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df47890a-73cc-4c48-84f2-a5c5fe217d35Cited by top-tier papers6
- Efficient Multivariate Time Series Forecasting via Calibrated Language Models with Privileged Knowledge DistillationChenxi Liu, Hao Miao, Qianxiong Xu, Shaowen Zhou et al.ICDE 2025 · 15 citations
- AimTS: Augmented Series and Image Contrastive Learning for Time Series ClassificationYuxuan Chen, Shanshan Huang, Yunyao Cheng, Peng Chen et al.ICDE 2025 · 5 citations
- Accurate and Efficient Multivariate Time Series Forecasting via Offline ClusteringYiming Niu, Jinliang Deng, Lulu Zhang, Zimu Zhou et al.ICDE 2025 · 4 citations
- DIFFODE: Neural ODE with Differentiable Hidden State for Irregular Time Series AnalysisYudong Zhang, Xu Wang, Xuan Yu, Zhengyang Zhou et al.ICDE 2025 · 3 citations
- From Teacher Pathways to Invariant Manifolds: Consensus Subspace Distillation for TSFMsZexing Zhang, Tianyang Lei, Jichao Li, Yang KeweiICML 2026
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
Related papers
- Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series ForecastingZezhi Shao, Zhao Zhang, Fei Wang, Yongjun XuKDD 2022 · 260 citations
- Towards a General Time Series Forecasting Model with Unified Representation and Adaptive TransferYihang Wang, Yuying Qiu, Peng Chen, Kai Zhao et al.ICML 2025
- GTM: A General Time-series Model for Enhanced Representation Learning of Time-Series dataCheng He, Xu Huang, Gangwei Jiang, Zhaoyi Li et al.ICLR 2026 · 4 citations
- Language Pre-training Guided Masking Representation Learning for Time Series ClassificationLiaoyuan Tang, Zheng Wang, Jie Wang, Guanxiong He et al.AAAI 2025 · 1 citation
- Time Without Time: Pseudo-Temporal Representation for Space-Time Super-ResolutionHee Min Choi, Hyoa Kang, Suji Kim, Dokwan Oh et al.CVPR 2026
