Enhancing Masked Time-Series Modeling via Dropping Patches
Tianyu Qiu, Yi Xie, Hao Niu, Yun Xiong, Xiaofeng Gao
Abstract
This paper explores how to enhance existing masked time-series modeling by randomly dropping sub-sequence level patches of time series. On this basis, a simple yet effective method named DropPatch is proposed, which has two remarkable advantages: 1) It improves the pre-training efficiency by a square-level advantage; 2) It provides additional advantages for modeling in scenarios such as in-domain, cross-domain, few-shot learning and cold start. This paper conducts comprehensive experiments to verify the effectiveness of the method and analyze its internal mechanism. Empirically, DropPatch strengthens the attention mechanism, reduces information redundancy and serves as an efficient means of data augmentation. Theoretically, it is proved that DropPatch slows down the rate at which the Transformer representations collapse into the rank-1 linear subspace by randomly dropping patches, thus optimizing the quality of the learned representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b52f117f-77ba-40c2-ac2c-2712641cafceCited by top-tier papers1
Ask how each one uses itBuilds on25
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
Related papers
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- Learning to Embed Time Series Patches IndependentlySeunghan Lee, Taeyoung Park, Kibok LeeICLR 2024 · 57 citations
- GTM: A General Time-series Model for Enhanced Representation Learning of Time-Series dataCheng He, Xu Huang, Gangwei Jiang, Zhaoyi Li et al.ICLR 2026 · 4 citations
- TimeSiam: A Pre-Training Framework for Siamese Time-Series ModelingJiaxiang Dong, Haixu Wu, Yuxuan Wang, Yunzhong Qiu et al.ICML 2024 · 23 citations
- Revisiting Token Dropping Strategy in Efficient BERT PretrainingQihuang Zhong, Liang Ding, Juhua Liu, Xuebo Liu et al.ACL 2023 · 5 citations
