Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
Annan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang, Abdul Fatir Ansari, Oleksandr Shchur, Christos Faloutsos, Andrew Gordon Wilson, Michael W. Mahoney, Bernie Wang
摘要
Transformers are widely used across data modalities, and yet the principles distilled from text models often transfer imperfectly to models trained to other modalities. In this paper, we analyze Transformers through the lens of rank structure. Our focus is on the time series setting, where the structural properties of the data differ remarkably from those of text or vision. We show that time-series embeddings, unlike text or vision, exhibit sharply decaying singular value spectra: small patch sizes and smooth continuous mappings concentrate the data into low-rank subspaces. From this, we prove that the associated projections admit accurate low-rank approximations, and that attention layers become compressible in proportion to the decay of the embedding spectrum. We introduce the concept of flow-of-ranks, a phenomenon by which nonlinear mixing across depth inflates the rank, explaining why early layers are most amenable to compression and why ranks grow with depth. Guided by these theoretical and empirical results, we use these insights to compress Chronos, a large time series foundation model, achieving a reduction of in inference time and in memory, without loss of accuracy. Our findings provide principled guidance for allocating width, depth, and heads in time series foundation models, and for exploiting their inherent compressibility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Understanding the Implicit Biases of Design Choices for Time Series Foundation ModelsAnnan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang 等ICLR 2026 · 被引用 11 次
- Zero-shot Forecasting by Simulation AloneBoris N. Oreshkin, Mayank Jauhari, Ravi Kiran Selvam, Malcolm Wolff 等ICLR 2026 · 被引用 4 次
- Universal Redundancies in Time Series Foundation ModelsAnthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai, William GilpinICML 2026 · 被引用 2 次
- Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?Xinyu Li, Kexi Chen, Jiajie Shen, Ying Zheng 等ACL 2026
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- N-BEATS: Neural basis expansion analysis for interpretable time series forecastingBoris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua BengioICLR 2020 · 被引用 1,550 次
相关 Paper
- Efficient Time Series Processing for Transformers and State-Space Models through Token MergingLeon Götz, Marcel Kollovieh, Stephan Günnemann, Leo SchwinnICML 2025
- Which transformer architecture fits my data? A vocabulary bottleneck in self-attentionNoam Wies, Yoav Levine, Daniel Jannai, Amnon ShashuaICML 2021 · 被引用 22 次
- Exploring Representations and Interventions in Time Series Foundation ModelsMichal Wilinski, Mononito Goswami, Willa Potosnak, Nina Zukowska 等ICML 2025
- LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse ApproximationYixiao Li, Yifan Yu, Qingru Zhang, Chen Liang 等ICML 2023 · 被引用 125 次
- Enhancing Foundation Models for Time Series Forecasting via Wavelet-based TokenizationLuca Masserano, Abdul Fatir Ansari, Boran Han, Xiyuan Zhang 等ICML 2025
