TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture Distillation
Juntong Ni, Zewen Liu, Shiyu Wang, Ming Jin, Wei Jin
摘要
Transformer-based and CNN-based methods demonstrate strong performance in long-term time series forecasting. However, their high computational and storage requirements can hinder large-scale deployment. To address this limitation, we propose integrating lightweight MLP with advanced architectures using knowledge distillation (KD). Our preliminary study reveals different models can capture complementary patterns, particularly multi-scale and multiperiod patterns in the temporal and frequency domains. Based on this observation, we introduce TimeDistill, a cross-architecture KD framework that transfers these patterns from teacher models (e.g., Transformers, CNNs) to MLP. Additionally, we provide a theoretical analysis, demonstrating that our KD approach can be interpreted as a specialized form of mixup data augmentation. TimeDistill improves MLP performance by up to 18.6%, surpassing teacher models on eight datasets. It also achieves up to 7× faster inference and requires 130× fewer parameters. Furthermore, we conduct extensive evaluations to highlight the versatility and effectiveness of TimeDistill. The code is available at Github Code Repo. CCS Concepts • Information systems → Temporal data; • Mathematics of computing → Time series analysis; • Computing methodologies → Neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement LearningJuntong Ni, Shiyu Wang, Qi He, Ming Jin 等ACL 2026 · 被引用 8 次
- TimeRecipe: A Time-Series Forecasting Recipe via Benchmarking Module Level EffectivenessZhiyuan Zhao, Juntong Ni, Shangqing Xu, Haoxin Liu 等ICLR 2026 · 被引用 7 次
- FusAD: Time-Frequency Fusion with Adaptive Denoising for General Time Series AnalysisDa Zhang, Bingyu Li, Zhiyuan Zhao, Feiping Nie 等ICDE 2026 · 被引用 4 次
- From Teacher Pathways to Invariant Manifolds: Consensus Subspace Distillation for TSFMsZexing Zhang, Tianyang Lei, Jichao Li, Yang KeweiICML 2026
它引用的顶会 Paper20
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 被引用 3,619 次
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang 等ICML 2022 · 被引用 2,912 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
相关 Paper
- Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal ForecastingYuqi Li, Chuanguang Yang, Hansheng Zeng, Zeyu Dong 等ICCV 2025 · 被引用 23 次
- Beyond Point Predictions: Manifold Expansion and Dual Alignment for Robust Time Series DistillationJunyao Hong, Zesheng Lai, Xinyi Xiao, Suyang Zhou 等ICML 2026
- Adaptive Multi-Scale Decomposition Framework for Time Series ForecastingYifan Hu, Peiyuan Liu, Peng Zhu, Dawei Cheng 等AAAI 2025 · 被引用 60 次
- TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series ForecastingVijay Ekambaram, Arindam Jati, Nam Nguyen, Phanwadee Sinthong 等KDD 2023 · 被引用 221 次
- Harmonic Dataset Distillation for Time Series ForecastingSeungha Hong, Sanghwan Jang, Wonbin Kweon, Suyeon Kim 等AAAI 2026
