Efficient Multivariate Time Series Forecasting via Calibrated Language Models with Privileged Knowledge Distillation
Chenxi Liu, Hao Miao, Qianxiong Xu, Shaowen Zhou, Cheng Long, Yan Zhao, Ziyue Li, Rui Zhao
Abstract
Multivariate time series forecasting (MTSF) endeavors to predict future observations given historical data, playing a crucial role in time series data management systems. With advancements in large language models (LLMs), recent studies employ textual prompt tuning to infuse the knowledge of LLMs into MTSF. However, the deployment of LLMs often suffers from low efficiency during the inference phase. To address this problem, we introduce TimeKD, an efficient MTSF framework that leverages the calibrated language models and privileged knowledge distillation. TimeKD aims to generate high-quality future representations from the proposed cross-modality teacher model and cultivate an effective student model. The cross-modality teacher model adopts calibrated language models (CLMs) with ground truth prompts, motivated by the paradigm of Learning Under Privileged Information (LUPI). In addition, we design a subtractive cross attention (SCA) mechanism to refine these representations. To cultivate an effective student model, we propose an innovative privileged knowledge distillation (PKD) mechanism including correlation and feature distillation. PKD enables the student to replicate the teacher's behavior while minimizing their output discrepancy. Extensive experiments on real data offer insight into the effectiveness, efficiency, and scalability of the proposed TimeKD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a09ff06-bb2d-4f3b-84cd-378a8167505bCited by top-tier papers24
- DBLoss: Decomposition-based Loss Function for Time Series ForecastingXiangfei Qiu, Xingjian Wu, Hanyin Cheng, Xvyuan Liu et al.NeurIPS 2025 · 61 citations
- Less but More: Linear Adaptive Graph Learning Empowering Spatiotemporal ForecastingJiaming Ma, Binwu Wang, Guanjun Wang, Kuo Yang et al.NeurIPS 2025 · 23 citations
- STRAP: Spatio-Temporal Pattern Retrieval for Out-of-Distribution GeneralizationHaoyu Zhang, Wentao Zhang, Hao Miao, Xinke Jiang et al.NeurIPS 2025 · 12 citations
- DisMS-TS: Eliminating Redundant Multi-scale Features for Time Series ClassificationZhipeng Liu, Peibo Duan, Binwu Wang, Xuan Tang et al.ACM MM 2025 · 5 citations
- SPOT-Trip: Dual-Preference Driven Out-of-Town Trip RecommendationYinghui Liu, Hao Miao, Guojiang Shen, Yan Zhao et al.NeurIPS 2025 · 5 citations
Builds on34
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun et al.NeurIPS 2023 · 1,178 citations
Related papers
- TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality AlignmentChenxi Liu, Qianxiong Xu, Hao Miao, Sun Yang et al.AAAI 2025 · 141 citations
- M3Time: LLM-Enhanced Multi-Modal, Multi-Scale, and Multi-Frequency Multivariate Time Series ForecastingShuning Jia, Baijun Song, Canming Ye, Chun YuanAAAI 2026 · 1 citation
- TimeMRA: LLM-Empowered Time Series Forecasting via Multi-Scale Retrieval-Augmented RepresentationsZongjiang Shang, Chengxi Jin, Binqing Wu, Dongliang Cui et al.ICML 2026
- Self-Improving Teacher Cultivates Better Student: Distillation Calibration for Multimodal Large Language ModelsXinwei Li, Li Lin, Shuai Wang, Chen QianSIGIR 2024 · 4 citations
- CALF: Aligning LLMs for Time Series Forecasting via Cross-modal Fine-TuningPeiyuan Liu, Hang Guo, Tao Dai, Naiqi Li et al.AAAI 2025 · 117 citations
