Provable Length Generalization in Sequence Prediction via Spectral Filtering
Annie Marsden, Evan Dogariu, Naman Agarwal, Xinyi Chen, Daniel Suo, Elad Hazan
Abstract
We consider the problem of length generalization in sequence prediction. We define a new metric of performance in this setting -the Asymmetric-Regret-which measures regret against a benchmark predictor with longer context length than available to the learner. We continue by studying this concept through the lens of the spectral filtering algorithm. We present a gradient-based learning algorithm that provably achieves length generalization for linear dynamical systems. We conclude with proof-of-concept experiments which are consistent with our theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Transformers Provably Learn Chain-of-Thought Reasoning with Length GeneralizationYu Huang, Zixin Wen, Aarti Singh, Yuejie Chi et al.NeurIPS 2025 · 22 citations
- SpectraLDS: Provable Distillation for Linear Dynamical SystemsDevan Shah, Shlomo Fortgang, Sofiia Druchyna, Elad HazanNeurIPS 2025 · 1 citation
- Next-Token Prediction and Regret MinimizationMehryar Mohri, Clayton Sanford, Jon Schneider, Kiran Vodrahalli et al.ICML 2026
- The Role of Sparsity for Length Generalization in LLMsNoah Golowich, Samy Jelassi, David Brandfonbrener, Sham M. Kakade et al.ICML 2025
- Non-Asymptotic Length GeneralizationThomas Chen, Tengyu Ma, Zhiyuan LiICML 2025
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
Related papers
- Length Generalization via Auxiliary TasksPranjal Awasthi, Anupam Gupta, Ravi KumarNeurIPS 2025
- Alignment-Sensitive Minimax Rates for Spectral Algorithms with Learned KernelsDongming Huang, Zhifan Li, Yicheng Li, Qian LinICML 2026
- Universal Learning of Nonlinear DynamicsEvan Dogariu, Anand Brahmbhatt, Elad HazanICML 2026 · 5 citations
- Long-Short Alignment for Effective Long-Context Modeling in LLMsTianqi Du, Haotian Huang, Yifei Wang, Yisen WangICML 2025
- SLIP: Learning to predict in unknown dynamical systems with long-term memoryParia Rashidinejad, Jiantao Jiao, Stuart RussellNeurIPS 2020 · 16 citations
