Learning to Embed Time Series Patches Independently
Seunghan Lee, Taeyoung Park, Kibok Lee
Abstract
Masked time series modeling has recently gained much attention as a self-supervised representation learning strategy for time series. Inspired by masked image modeling in computer vision, recent works first patchify and partially mask out time series, and then train Transformers to capture the dependencies between patches by predicting masked patches from unmasked patches. However, we argue that capturing such patch dependencies might not be an optimal strategy for time series representation learning; rather, learning to embed patches independently results in better time series representations. Specifically, we propose to use 1) the simple patch reconstruction task, which autoencode each patch without looking at other patches, and 2) the simple patch-wise MLP that embeds each patch independently. In addition, we introduce complementary contrastive learning to hierarchically capture adjacent time series information efficiently. Our proposed method improves time series forecasting and classification performance compared to state-of-the-art Transformer-based models, while it is more efficient in terms of the number of parameters and training/inference time. Code is available at this repository: https://github.com/seunghan96/pits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a84181ea-ddd0-48c2-bde8-348380bcf152Cited by top-tier papers10
- Are Language Models Actually Useful for Time Series Forecasting?Mingtian Tan, Mike A. Merrill, Vinayak Gupta, Tim Althoff et al.NeurIPS 2024 · 326 citations
- UniTS: A Unified Multi-Task Time Series ModelShanghua Gao, Teddy Koker, Owen Queen, Tom Hartvigsen et al.NeurIPS 2024 · 159 citations
- Rethinking Channel Dependence for Multivariate Time Series Forecasting: Learning from Leading IndicatorsLifan Zhao, Yanyan ShenICLR 2024 · 50 citations
- Unlocking the Power of Patch: Patch-Based MLP for Long-Term Time Series ForecastingPeiwang Tang, Weitai ZhangAAAI 2025 · 42 citations
- Multi-Patch Prediction: Adapting Language Models for Time Series Representation LearningYuxuan Bian, Xuan Ju, Jiangtong Li, Zhijian Xu et al.ICML 2024 · 13 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
Related papers
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- PPT: Patch Order Do Matters In Time Series Pretext TaskJaeho Kim, Kwangryeol Park, Sukmin Yun, Seulki LeeICLR 2025
- Language Pre-training Guided Masking Representation Learning for Time Series ClassificationLiaoyuan Tang, Zheng Wang, Jie Wang, Guanxiong He et al.AAAI 2025 · 1 citation
- Mantis: Lightweight Foundation Model for Time Series ClassificationVasilii Feofanov, Songkang Wen, Shifeng Xie, Simon Roschmann et al.ICML 2026
- PatchCLE: Breaking the Linear Representation Bottleneck in Time Series Forecasting via Soft Contrastive LearningMengSen Wu, Haochen Shi, Shengdong Du, Jie Hu et al.KDD 2026
