Two-Layer Linear Auto-Regressive Models Estimate Latent States
Yahya Sattar, Sunmook Choi, Leo Maynard-Zhang, Yassir Jedra, Maryam Fazel, Sarah Dean
Abstract
Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models learn latent representations remains an open theoretical question. In this work, we demonstrate that when trained by empirical risk minimization on data from partially observed linear dynamical systems, two-layer linear auto-regressive models naturally learn to approximate Kalman filtering. In particular, we show that the learned hidden representation coincides, up to a similarity transformation, with the state estimates produced by the optimal (Kalman) filter, even though the model has no explicit knowledge of the underlying dynamics or state. The result follows from three main insights. First, we establish that the Kalman filter is well approximated by an auto-regressive model with bounded truncation error. Second, we show that despite non-convexity, the two-layer optimization landscape is benign, i.e., all stationary points are either strict saddles or global minima. Finally, as our main contributions, we provide finite-sample guarantees on prediction error, parameter estimation error, and latent state recovery. Numerical simulations support the theoretical results and demonstrate that the latent representations of auto-regressive models recover state estimates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf5bffc7-34b6-4e37-bb69-f93d26e1ca1aBuilds on7
- Logarithmic Regret Bound in Partially Observable Linear Dynamical SystemsSahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, Anima AnandkumarNeurIPS 2020 · 106 citations
- Chess as a Testbed for Language Model State TrackingShubham Toshniwal, Sam Wiseman, Karen Livescu, Kevin GimpelAAAI 2022 · 77 citations
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas et al.ICLR 2023 · 60 citations
- Learning with little mixingIngvar M. Ziemann, Stephen TuNeurIPS 2022 · 41 citations
- A New Approach to Learning Linear Dynamical SystemsAinesh Bakshi, Allen Liu, Ankur Moitra, Morris YauSTOC 2023 · 10 citations
Related papers
- Self-Supervised Inference in State-Space ModelsDavid Ruhe, Patrick ForréICLR 2022 · 8 citations
- On learning linear dynamical systems in context with attention layersMaria-Luiza Vladarean, Xuhui Zhang, Suvrit SraICLR 2026
- Why Linear Recurrent Memory Works in Partially Observable Reinforcement LearningYike Zhao, Onno Eberhard, Malek khammassi, Ali Sayed et al.ICML 2026 · 2 citations
- Extracting Latent State Representations with Linear Dynamics from Rich ObservationsAbraham Frandsen, Rong Ge, Holden LeeICML 2022 · 1 citation
- Identifying latent state transitions in non-linear dynamical systemsÇaglar Hizli, Çagatay Yildiz, Matthias Bethge, S. T. John et al.ICLR 2025
