Byte Pair Encoding for Efficient Time Series Forecasting
Leon Götz, Marcel Kollovieh, Stephan Günnemann, Leo Schwinn
Abstract
Existing time series tokenization methods predominantly encode a constant number of samples into individual tokens. This inflexible approach can generate excessive tokens for even simple patterns like extended constant values, resulting in substantial computational overhead. Inspired by the success of byte pair encoding, we propose the first pattern-centric tokenization scheme for time series analysis. Based on a discrete vocabulary of frequent motifs, our method merges samples with underlying patterns into tokens, compressing time series adaptively. Exploiting our finite set of motifs and the continuous properties of time series, we further introduce conditional decoding as a lightweight yet powerful post-hoc optimization method, which requires no gradient computation and adds no computational overhead. On recent time series foundation models, our motifbased tokenization improves forecasting performance by 40 % and boosts efficiency by 2314 % on average. Conditional decoding further reduces MSE by up to 48 %. In an extensive analysis, we demonstrate the adaptiveness of our tokenization to diverse temporal patterns, its generalization to unseen data, and its meaningful token representations capturing distinct time series properties, including statistical moments and trends.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b67f7ff-5e6d-4636-bc80-35f89d23fcabCited by top-tier papers3
- OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series DataPatrick Langer, Thomas Kaar, Max Rosenblattl, Maxwell A. Xu et al.ICML 2026 · 22 citations
- Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and CompressibilityAnnan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang et al.ICLR 2026 · 9 citations
- Dywave: Event-Aligned Dynamic Tokenization for Heterogeneous IoT Sensing SignalsTomoyoshi Kimura, Denizhan Kara, Jinyang Li, Hongjue Zhao et al.ICML 2026
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- Non-stationary Transformers: Exploring the Stationarity in Time Series ForecastingYong Liu, Haixu Wu, Jianmin Wang, Mingsheng LongNeurIPS 2022 · 1,080 citations
Related papers
- Enhancing Foundation Models for Time Series Forecasting via Wavelet-based TokenizationLuca Masserano, Abdul Fatir Ansari, Boran Han, Xiyuan Zhang et al.ICML 2025
- LightGTS: A Lightweight General Time Series Forecasting ModelYihang Wang, Yuying Qiu, Peng Chen, Yang Shu et al.ICML 2025
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 6 citations
- Mantis: Lightweight Foundation Model for Time Series ClassificationVasilii Feofanov, Songkang Wen, Shifeng Xie, Simon Roschmann et al.ICML 2026
