Dynamic Relational Priming Improves Transformer in Multivariate Time Series
Hunjae Lee, Corey Clark
摘要
Standard attention in transformers employ static token representations that remain unchanged across all pair-wise computations in each layer. This limits their representational alignment with the potentially diverse dynamics of each tokenpair interaction. While they excel in domains with relatively homogeneous relationships, standard attention may be inadequate in capturing heterogeneous inter-channel dependencies of multivariate time series (MTS) data where different channelpair interactions within a single system may be governed by entirely different physical laws or temporal dynamics. To better align the attention mechanism for such domain phenomena, we propose attention with dynamic relational priming (prime attention). Prime attention modulates token representations for each token-pair, optimizing each pair-wise interaction for that specific relationship. Results demonstrate that prime attention consistently outperforms standard attention across benchmarks, achieving up to 6.5% improvement in forecasting accuracy. In addition, prime attention achieves comparable performance using up to 40% less sequence length compared to standard attention, demonstrating its superior relational modeling capabilities and potential for data efficiency. Code is available at https://github.com/ timlee0131/Prime-Attention .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 被引用 3,619 次
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu 等NeurIPS 2022 · 被引用 1,216 次
相关 Paper
- Sequence Complementor: Complementing Transformers for Time Series Forecasting with Learnable SequencesXiwen Chen, Peijie Qiu, Wenhui Zhu, Huayu Li 等AAAI 2025 · 被引用 4 次
- GraFT: Infusing Pre-trained Transformers with Relational Structure for Time Series ForecastingYuqi Yuan, Xiong Luo, Qiaojuan Peng, Wenbing ZhaoAAAI 2026
- Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?Xinyu Li, Kexi Chen, Jiajie Shen, Ying Zheng 等ACL 2026
- PHAT: Modeling Period Heterogeneity for Multivariate Time Series ForecastingJiaming Ma, Qihe Huang, Haofeng Ma, Guanjun Wang 等ICLR 2026 · 被引用 6 次
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 被引用 536 次
