Scalable Pre-Training of Compact Urban Spatio-Temporal Predictive Models on Large-Scale Multi-Domain Data
Jindong Han, Hao Wang, Hui Xiong, Hao Liu
Abstract
Spatio-Temporal Prediction (STP) is crucial for various smart city applications, such as traffic management and resource allocation. However, training samples can be scarce in data-constrained scenarios, which often degrades the predictive capability of existing deep STP models. Although recent STP foundation models excel in few-shot and zero-shot learning through extensive pre-training on large-scale, multi-domain spatio-temporal data, they often rely on large parameter scale to achieve enhanced performance, resulting in high computational demands that hinder practical deployment. In response, we develop CompactST, an efficient, compact, and versatile pre-trained model for STP in data-scarce settings. Recognizing the complexities posed by large-scale, heterogeneous pre-training datasets, CompactST integrates three specialized components: (1) a mixture-of-normalizers module to address domain and spatial heterogeneity, (2) a multi-scale spatio-temporal mixer that captures diverse patterns from datasets with varying spatio-temporal resolutions, and (3) an adaptive dataset-oriented tuning module that transfers the handling of dataset-specific parameters from pre-training to fine-tuning stage. These tailored designs enable CompactST to maximize generalizability across diverse datasets while maintaining a compact model size ( i.e. , only 300K parameters). To validate its effectiveness, we pre-train CompactST on a substantial corpus of public spatio-temporal datasets spanning over 10 domains and encompassing 300 million data points. Extensive experimental results on ten real-world datasets demonstrate CompactST's significantly improved prediction accuracy and efficiency in data-scarce scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 017b3be8-cda1-427d-92e5-3ab6348dc14eBuilds on34
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural NetworksZonghan Wu, Shirui Pan, Guodong Long, Jing Jiang et al.KDD 2020 · 1,738 citations
Related papers
- UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal PredictionYuan Yuan, Jingtao Ding, Jie Feng, Depeng Jin et al.KDD 2024 · 75 citations
- Spatio-Temporal Few-Shot Learning via Diffusive Neural Network GenerationYuan Yuan, Chenyang Shao, Jingtao Ding, Depeng Jin et al.ICLR 2024 · 34 citations
- Damba-ST: Domain-Adaptive Mamba for Efficient Urban Spatio-Temporal PredictionRui An, Yifeng Zhang, Ziran Liang, Wenqi Fan et al.ICDE 2026
- AutoST: Efficient Neural Architecture Search for Spatio-Temporal PredictionTing Li, Junbo Zhang, Kainan Bao, Yuxuan Liang et al.KDD 2020 · 86 citations
- CrossST: An Efficient Pre-Training Framework for Cross-District Pattern Generalization in Urban Spatio-Temporal ForecastingAoyu Liu, Yaying ZhangICDE 2025 · 3 citations
