Robust Inter-Series Dependency Modeling for Time Series Forecasting via Information-Theoretic Alignment
Wuqing Yu, Weichen Guo, Jian Zhou, Shuyu Luo, Jiacai Zhang
Abstract
While iTransformer pioneered general inter-variate dependency (IVD) modeling in Transformers for multivariate time series forecasting (MTSF), subsequent research on such universal paradigms has been surprisingly scarce. Through comprehensive analysis, we identify a critical structural inconsistency in Variate Transformers: typically capturing inter-variate dependencies via shallow self-attention layers while neglecting the critical requirement for deep-layer IVD modeling, which causes spurious correlations modeling and difficulties in model optimization. To address these limitations, we propose CGTFra, as a general framework for consistent IVD modeling. Specifically, we reconsider existing timestamp-based modeling and introduce a frequency-domain masking and resampling method for periodicity preservation, which serves as a general strategy for input feature enhancement. Additionally, CGTFra promotes consistent IVD modeling from two perspectives. Initially, a dynamic graph learning framework is integrated into Transformers to explicitly model IVD in deep network layer. Furthermore, grounded in the Information Bottleneck principle, we further propose a consistency-constrained alignment to learn more robust IVD and temporal feature representations. These three core design philosophies of CGTFra can be integrated into any existing Variate Transformer-based framework, and CGTFra achieves superior predictive performance across 13 long- and short-term datasets with high computational efficiency and desirable interpretability. Code is available at https://github.com/05Pikachu24/Consistent-CGTFra.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a475e9a-b4b5-43b3-bd01-e8bd3ffce669Builds on7
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu et al.ICLR 2024 · 1,703 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- Pure Transformers are Powerful Graph LearnersJinwoo Kim, Dat Nguyen, Seonwoo Min, Sungjun Cho et al.NeurIPS 2022 · 311 citations
- TSLANet: Rethinking Transformers for Time Series Representation LearningEmadeldeen Eldele, Mohamed Ragab, Zhenghua Chen, Min Wu et al.ICML 2024 · 159 citations
Related papers
- GraFT: Infusing Pre-trained Transformers with Relational Structure for Time Series ForecastingYuqi Yuan, Xiong Luo, Qiaojuan Peng, Wenbing ZhaoAAAI 2026
- Crossformer: Transformer Utilizing Cross-Dimension Dependency for Multivariate Time Series ForecastingYunhao Zhang, Junchi YanICLR 2023
- PESD-TSF: A Period-Aware and Explicit Structured Decomposition Framework for Long-Term Time Series ForecastingHua Wang, Xianhao Jiao, Fan ZhangICML 2026
- Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive ForecastingJiecheng Lu, Shihao YangICML 2025
- A Memory Guided Transformer for Time Series ForecastingYunyao Cheng, Chenjuan Guo, Bin Yang, Haomin Yu et al.VLDB 2025 · 1 citation
