Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
Annan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang, Abdul Fatir Ansari, Oleksandr Shchur, Christos Faloutsos, Andrew Gordon Wilson, Michael W. Mahoney, Bernie Wang
Abstract
Transformers are widely used across data modalities, and yet the principles distilled from text models often transfer imperfectly to models trained to other modalities. In this paper, we analyze Transformers through the lens of rank structure. Our focus is on the time series setting, where the structural properties of the data differ remarkably from those of text or vision. We show that time-series embeddings, unlike text or vision, exhibit sharply decaying singular value spectra: small patch sizes and smooth continuous mappings concentrate the data into low-rank subspaces. From this, we prove that the associated projections admit accurate low-rank approximations, and that attention layers become compressible in proportion to the decay of the embedding spectrum. We introduce the concept of flow-of-ranks, a phenomenon by which nonlinear mixing across depth inflates the rank, explaining why early layers are most amenable to compression and why ranks grow with depth. Guided by these theoretical and empirical results, we use these insights to compress Chronos, a large time series foundation model, achieving a reduction of in inference time and in memory, without loss of accuracy. Our findings provide principled guidance for allocating width, depth, and heads in time series foundation models, and for exploiting their inherent compressibility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2222cd3-4b99-4bd6-ac95-45c5a496d97fCited by top-tier papers4
- Understanding the Implicit Biases of Design Choices for Time Series Foundation ModelsAnnan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang et al.ICLR 2026 · 11 citations
- Zero-shot Forecasting by Simulation AloneBoris N. Oreshkin, Mayank Jauhari, Ravi Kiran Selvam, Malcolm Wolff et al.ICLR 2026 · 4 citations
- Universal Redundancies in Time Series Foundation ModelsAnthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai, William GilpinICML 2026 · 2 citations
- Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?Xinyu Li, Kexi Chen, Jiajie Shen, Ying Zheng et al.ACL 2026
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- N-BEATS: Neural basis expansion analysis for interpretable time series forecastingBoris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua BengioICLR 2020 · 1,550 citations
Related papers
- Efficient Time Series Processing for Transformers and State-Space Models through Token MergingLeon Götz, Marcel Kollovieh, Stephan Günnemann, Leo SchwinnICML 2025
- Which transformer architecture fits my data? A vocabulary bottleneck in self-attentionNoam Wies, Yoav Levine, Daniel Jannai, Amnon ShashuaICML 2021 · 22 citations
- Exploring Representations and Interventions in Time Series Foundation ModelsMichal Wilinski, Mononito Goswami, Willa Potosnak, Nina Zukowska et al.ICML 2025
- LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse ApproximationYixiao Li, Yifan Yu, Qingru Zhang, Chen Liang et al.ICML 2023 · 125 citations
- Enhancing Foundation Models for Time Series Forecasting via Wavelet-based TokenizationLuca Masserano, Abdul Fatir Ansari, Boran Han, Xiyuan Zhang et al.ICML 2025
