Less is More: Unlocking Specialization of Time Series Foundation Models via Structured Pruning
Lifan Zhao, Yanyan Shen, Zhaoyang Liu, Xue Wang, Jiaji Deng
Abstract
Scaling laws motivate the development of Time Series Foundation Models (TSFMs) that pre-train vast parameters and achieve remarkable zero-shot forecasting performance. Surprisingly, even after fine-tuning, TSFMs cannot consistently outperform smaller, specialized models trained on full-shot downstream data. A key question is how to realize effective adaptation of TSFMs for a target forecasting task. Through empirical studies on various TSFMs, the pre-trained models often exhibit inherent sparsity and redundancy in computation, suggesting that TSFMs have learned to activate task-relevant network substructures to accommodate diverse forecasting tasks. To preserve this valuable prior knowledge, we propose a structured pruning method to regularize the subsequent fine-tuning process by focusing it on a more relevant and compact parameter space. Extensive experiments on seven TSFMs and six benchmarks demonstrate that fine-tuning a smaller, pruned TSFM significantly improves forecasting performance compared to fine-tuning original models. This "prune-then-finetune" paradigm often enables TSFMs to achieve state-of-the-art performance and surpass strong specialized baselines. Source code is made publicly available at https://github.com/SJTU-DMTai/Prune-then-Finetune .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 439d3877-e581-4e16-a068-e80525676e1aCited by top-tier papers2
- TS-Memory: Plug-and-Play Memory for Time Series Foundation ModelsSisuo Lyu, Siru Zhong, Tiegang Chen, Weilin Ruan et al.KDD 2026
- Revealing Scaling Paradox in Large-scale Time Series Models: Implications for More Efficient and Accurate ForecastingXin Qiu, Junlong Tong, Yirong Sun, Yunpu Ma et al.ICML 2026
Builds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution ShiftTaesung Kim, Jinhee Kim, Yunwon Tae, Cheonbok Park et al.ICLR 2022 · 1,020 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
Related papers
- Multi-Scale Finetuning for Encoder-based Time Series Foundation ModelsZhongzheng Qiao, Chenghao Liu, Yiming Zhang, Ming Jin et al.NeurIPS 2025 · 17 citations
- Exploring Representations and Interventions in Time Series Foundation ModelsMichal Wilinski, Mononito Goswami, Willa Potosnak, Nina Zukowska et al.ICML 2025
- SEMPO: Lightweight Foundation Models for Time Series ForecastingHui He, Kun Yi, Yuanchi Ma, Qi Zhang et al.NeurIPS 2025 · 12 citations
- Meta-Learning Framework with Applications to Zero-Shot Time-Series ForecastingBoris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua BengioAAAI 2021 · 135 citations
- In-Context Fine-Tuning for Time-Series Foundation ModelsMatthew Faw, Rajat Sen, Yichen Zhou, Abhimanyu DasICML 2025
