Universal Sequence Preconditioning
Annie Marsden, Elad Hazan
Abstract
We study the problem of preconditioning in sequential prediction. From the theoretical lens of linear dynamical systems, we show that convolving the target sequence corresponds to applying a polynomial to the hidden transition matrix. Building on this insight, we propose a universal preconditioning method that convolves the target with coefficients from orthogonal polynomials such as Chebyshev or Legendre. We prove that this approach reduces regret for two distinct prediction algorithms and yields the first ever sublinear and hidden-dimension-independent regret bounds (up to logarithmic factors) that hold for systems with marginally stable and asymmetric transition matrices. Finally, extensive synthetic and realworld experiments show that this simple preconditioning strategy improves the performance of a diverse range of algorithms, including recurrent neural networks, and generalizes to signals beyond linear dynamical systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Universal Learning of Nonlinear DynamicsEvan Dogariu, Anand Brahmbhatt, Elad HazanICML 2026 · 5 citations
- SpectraLDS: Provable Distillation for Linear Dynamical SystemsDevan Shah, Shlomo Fortgang, Sofiia Druchyna, Elad HazanNeurIPS 2025 · 1 citation
Builds on8
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
Related papers
- Anamnesic Neural Differential Equations with Orthogonal Polynomial ProjectionsEdward De Brouwer, Rahul G. KrishnanICLR 2023 · 2 citations
- Polynomial Preconditioning for Gradient MethodsNikita Doikov, Anton RodomanovICML 2023 · 2 citations
- Provable Length Generalization in Sequence Prediction via Spectral FilteringAnnie Marsden, Evan Dogariu, Naman Agarwal, Xinyi Chen et al.ICML 2025
- On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization AnalysisZhong Li, Jiequn Han, Weinan E, Qianxiao LiICLR 2021 · 40 citations
- Chebyshev Policies and the Mountain Car Problem: Reinforcement Learning for Low-dimensional Control TasksStefan Huber, Hannes Unger, Georg Schäfer, Jakob RehrlICML 2026
