Baguan-TS: dual in-context learning model for time series forecasting with covariates
Linxiao Yang, Xue Jiang, Gezheng Xu, Tian Zhou, Min Yang, Zhaoyang Zhu, Linyuan Geng, Zhipeng Zeng, Qiming Chen, Xinyue Gu, Rong Jin, Liang Sun
摘要
Transformers enable in-context learning (ICL) for rapid, gradient-free adaptation in time series forecasting, yet most ICL-style approaches rely on tabularized, hand-crafted features, while end-toend sequence models lack inference-time adaptation. We bridge this gap with a unified framework, Baguan-TS, which integrates the raw-sequence representation learning with ICL, instantiated by a 3D Transformer that attends jointly over temporal, variable, and context axes. To make this high-capacity model practical, we tackle two key hurdles: (i) calibration and training stability, improved with a feature-agnostic, target-space retrieval-based local calibration; and (ii) output oversmoothing, mitigated via context-overfitting strategy. On public benchmark with covariates, Baguan-TS consistently outperforms established baselines, achieving the highest win rate and significant reductions in both point and probabilistic forecasting metrics. Further evaluations across diverse real-world energy datasets demonstrate its robustness, yielding substantial improvements.
• We develop a Y-space RBfcst local calibration module-feature-agnostic and episode-specific-that improves calibration, data efficiency, and scalability when training larger 3D Transformers. In our experiments, this module consistently improves overall forecasting accuracy and robustness under injected noise compared to training without retrieval.
• We introduce a context-overfitting strategy that explicitly balances sample denoising and sample selection, stabilizing in-context learning in high-capacity models. The strategy consistently lowers training loss and restores periodic spike reconstruction, mitigating oversmoothing without harming trend accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang 等ICML 2022 · 被引用 2,912 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
- N-BEATS: Neural basis expansion analysis for interpretable time series forecastingBoris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua BengioICLR 2020 · 被引用 1,550 次
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun 等NeurIPS 2023 · 被引用 1,178 次
相关 Paper
- In-context Time Series PredictorJiecheng Lu, Yan Sun, Shihao YangICLR 2025
- Enhancing Large Language Models for Time-Series Forecasting via Vector-Injected In-Context LearningJianqi Zhang, Jingyao Wang, Wenwen Qiang, Fanjiang Xu 等WWW 2026
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty 等KDD 2021 · 被引用 66 次
- One Step Closer to Ground Truth: A Multi-Scale Residual-Aware Representation Learning Pipeline for Predicting Time Series DataAmrijit Biswas, Mustafa Kamal, Robin Krambroeckers, Mirza M. Lutfe Elahi 等KDD 2026 · 被引用 1 次
- Timer-XL: Long-Context Transformers for Unified Time Series ForecastingYong Liu, Guo Qin, Xiangdong Huang, Jianmin Wang 等ICLR 2025
