Baguan-TS: dual in-context learning model for time series forecasting with covariates
Linxiao Yang, Xue Jiang, Gezheng Xu, Tian Zhou, Min Yang, Zhaoyang Zhu, Linyuan Geng, Zhipeng Zeng, Qiming Chen, Xinyue Gu, Rong Jin, Liang Sun
Abstract
Transformers enable in-context learning (ICL) for rapid, gradient-free adaptation in time series forecasting, yet most ICL-style approaches rely on tabularized, hand-crafted features, while end-toend sequence models lack inference-time adaptation. We bridge this gap with a unified framework, Baguan-TS, which integrates the raw-sequence representation learning with ICL, instantiated by a 3D Transformer that attends jointly over temporal, variable, and context axes. To make this high-capacity model practical, we tackle two key hurdles: (i) calibration and training stability, improved with a feature-agnostic, target-space retrieval-based local calibration; and (ii) output oversmoothing, mitigated via context-overfitting strategy. On public benchmark with covariates, Baguan-TS consistently outperforms established baselines, achieving the highest win rate and significant reductions in both point and probabilistic forecasting metrics. Further evaluations across diverse real-world energy datasets demonstrate its robustness, yielding substantial improvements.
• We develop a Y-space RBfcst local calibration module-feature-agnostic and episode-specific-that improves calibration, data efficiency, and scalability when training larger 3D Transformers. In our experiments, this module consistently improves overall forecasting accuracy and robustness under injected noise compared to training without retrieval.
• We introduce a context-overfitting strategy that explicitly balances sample denoising and sample selection, stabilizing in-context learning in high-capacity models. The strategy consistently lowers training loss and restores periodic spike reconstruction, mitigating oversmoothing without harming trend accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7436f7b1-8ec5-4ca6-8f58-5a88467a69eaBuilds on23
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- N-BEATS: Neural basis expansion analysis for interpretable time series forecastingBoris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua BengioICLR 2020 · 1,550 citations
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun et al.NeurIPS 2023 · 1,178 citations
Related papers
- In-context Time Series PredictorJiecheng Lu, Yan Sun, Shihao YangICLR 2025
- Enhancing Large Language Models for Time-Series Forecasting via Vector-Injected In-Context LearningJianqi Zhang, Jingyao Wang, Wenwen Qiang, Fanjiang Xu et al.WWW 2026
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty et al.KDD 2021 · 66 citations
- One Step Closer to Ground Truth: A Multi-Scale Residual-Aware Representation Learning Pipeline for Predicting Time Series DataAmrijit Biswas, Mustafa Kamal, Robin Krambroeckers, Mirza M. Lutfe Elahi et al.KDD 2026 · 1 citation
- Timer-XL: Long-Context Transformers for Unified Time Series ForecastingYong Liu, Guo Qin, Xiangdong Huang, Jianmin Wang et al.ICLR 2025
