TimesBERT: A BERT-Style Foundation Model for Time Series Understanding
Haoran Zhang, Yong Liu, Yunzhong Qiu, Haixuan Liu, Zhongyi Pei, Jianmin Wang, Mingsheng Long
Abstract
Time series analysis is crucial in diverse scenarios. Beyond forecasting, considerable real-world tasks are categorized into classification, imputation, and anomaly detection, underscoring different capabilities termed time series understanding in this paper. While GPT-style models have been positioned as foundation models for time series forecasting, the BERT-style architecture, which has made significant advances in natural language understanding, has not been fully unlocked for time series understanding, possibly attributed to the undesirable dropout of essential elements of BERT. In this paper, inspired by the shared multi-granularity structure between multivariate time series and multisentence documents, we design TimesBERT to learn generic representations of time series including temporal patterns and variate-centric characteristics. In addition to a natural adaptation of masked modeling, we propose a parallel task of functional token prediction to embody vital multi-granularity structures. Our model is pre-trained on 260 billion time points across diverse domains. Leveraging multi-granularity representations, TimesBERT achieves state-of-the-art performance across four typical downstream understanding tasks, outperforming task-specific models and language pre-trained backbones, positioning it as a versatile foundation model for time series understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3eb5099a-3e4e-4fcf-915f-63e208c0214eCited by top-tier papers2
- CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic DataShifeng Xie, Vasilii Feofanov, Jianfeng Zhang, Themis Palpanas et al.ICLR 2026 · 15 citations
- FeDaL: Federated Dataset Learning for General Time Series Foundation ModelsShengchao Chen, Guodong Long, Michael Blumenstein, Jing JiangICLR 2026 · 11 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Timer: Generative Pre-trained Transformers Are Large Time Series ModelsYong Liu, Haoran Zhang, Chenyu Li, Xiangdong Huang et al.ICML 2024 · 188 citations
- GTM: A General Time-series Model for Enhanced Representation Learning of Time-Series dataCheng He, Xu Huang, Gangwei Jiang, Zhaoyi Li et al.ICLR 2026 · 4 citations
- UniTS: A Unified Multi-Task Time Series ModelShanghua Gao, Teddy Koker, Owen Queen, Tom Hartvigsen et al.NeurIPS 2024 · 159 citations
- TimesNet: Temporal 2D-Variation Modeling for General Time Series AnalysisHaixu Wu, Tengge Hu, Yong Liu, Hang Zhou et al.ICLR 2023 · 423 citations
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty et al.KDD 2021 · 66 citations
