Selective Learning for Deep Time Series Forecasting
Yisong Fu, Zezhi Shao, Chengqing Yu, Yujie Li, Zhulin An, Qi (Cheems) Wang, Yongjun Xu, Fei Wang
Abstract
Benefiting from high capacity for capturing complex temporal patterns, deep learning (DL) has significantly advanced time series forecasting (TSF). However, deep models tend to suffer from severe overfitting due to the inherent vulnerability of time series to noise and anomalies. The prevailing DL paradigm uniformly optimizes all timesteps through the MSE loss and learns those uncertain and anomalous timesteps without difference, ultimately resulting in overfitting. To address this, we propose a novel selective learning strategy for deep TSF. Specifically, selective learning screens a subset of the whole timesteps to calculate the MSE loss in optimization, guiding the model to focus on generalizable timesteps while disregarding non-generalizable ones. Our framework introduces a dual-mask mechanism to target timesteps: (1) an uncertainty mask leveraging residual entropy to filter uncertain timesteps, and (2) an anomaly mask employing residual lower bound estimation to exclude anomalous timesteps. Extensive experiments across eight real-world datasets demonstrate that selective learning can significantly improve the predictive performance for typical state-of-the-art deep models, including 37.4% MSE reduction for Informer, 8.4% for TimesNet, and 6.5% for iTransformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25cb2811-7bf4-488e-a430-e8bab0cb224dCited by top-tier papers3
- DropoutTS: Sample-Adaptive Dropout for Robust Time Series ForecastingSiru Zhong, Yiqiu Liu, Zhiqing Cui, Zezhi Shao et al.ICML 2026
- APT: Affine Prototype-Timestamp for Time Series Forecasting Under Distribution ShiftYujie Li, Zezhi Shao, Chengqing Yu, Yisong Fu et al.AAAI 2026
- Zeus: Towards Tuning-Free Foundation Model for Time Series AnalysisYisong Fu, Zezhi Shao, Chengqing Yu, Yujie Li et al.ICML 2026
Builds on53
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural NetworksZonghan Wu, Shirui Pan, Guodong Long, Jing Jiang et al.KDD 2020 · 1,738 citations
Related papers
- Abstain Mask Retain Core: Time Series Prediction by Adaptive Masking Loss with Representation ConsistencyRenzhao Liang, Sizhe Xu, Chenggang Xie, Jingru Chen et al.NeurIPS 2025 · 2 citations
- From Observations to States: Latent Time Series ForecastingJie Yang, Yifan Hu, Yuante Li, Kexin Zhang et al.ICML 2026 · 3 citations
- Amortized Predictability-aware Training Framework for Time Series Forecasting and ClassificationXu Zhang, Peng Wang, Yichen Li, Wei WangWWW 2026
- Hierarchical Classification Auxiliary Network for Time Series ForecastingYanru Sun, Zongxia Xie, Dongyue Chen, Emadeldeen Eldele et al.AAAI 2025 · 28 citations
- Robust Inter-Series Dependency Modeling for Time Series Forecasting via Information-Theoretic AlignmentWuqing Yu, Weichen Guo, Jian Zhou, Shuyu Luo et al.ICML 2026
